> For the complete documentation index, see [llms.txt](https://docs.bito.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.bito.ai/governor/set-up-bito-governor-self-hosted-with-docker.md).

# Set up Bito Governor (self-hosted with Docker)

Run Bito Governor inside your own network. Install it on a single host with Docker, configure your workspace, and point your coding tools at it, with all traffic staying local.

Install Bito Governor on one machine, sign in, connect your AI coding tools, and operate it. This is the **self-hosted, single-node** path (gateway + database on one box, via Docker). No source code or build tools needed — just Docker.

{% hint style="info" %}
**What Bito Governor is:** a gateway that sits between your AI coding tools (Claude Code, Cursor, Codex, …) and the LLM providers. You point your tools at *one* gateway URL + key; the gateway routes to the right provider, and meters usage and cost per workspace.
{% endhint %}

## Prerequisites

You need the following:

* **Docker** — Docker Desktop (macOS/Windows) or Docker Engine + Compose (Linux). That's the only prerequisite; the gateway and its database run as containers.
* **A free port** — the gateway serves on **8788** by default (set `PORT` to change it).
* **An API key for each LLM provider you use,** such as Anthropic, OpenAI, Groq, or Google. Governor sends requests to your own provider accounts.
* **The model names your coding tools request,** such as `claude-opus-4-8`. Model names are case sensitive.
* **AI Architect MCP URL and access token,** if you enable AI Architect. Both are available in your [Bito account](https://alpha.bito.ai/).

***

## Step 1: Install

Run the one-liner for your OS. It pulls the Bito Governor image, brings up the stack, generates its secrets, signs you in, and prints everything you need.

**Linux / macOS**

```bash
curl -fsSL https://gwinstall.bito.ai/latest/install-standalone.sh | bash
```

**Windows (PowerShell)**

```powershell
irm https://gwinstall.bito.ai/latest/install-standalone.ps1 | iex
```

Everything installs under a home directory — **`~/.bito-gateway`** (Windows: `%USERPROFILE%\.bito-gateway`). Set **`BITO_GW_HOME`** to install elsewhere.

{% hint style="info" %}
Prefer a native Linux service (systemd) instead of Docker? See `docs/dev/DEPLOY_RUNBOOK.md` in your bito-gateway installation folder.
{% endhint %}

## Step 2: What you get

When it finishes you'll see a **"ready" block** like this — **save what it tells you to**:

```
════════════════════════════════════════════════════════════════
  ✓ bito-gateway is ready
════════════════════════════════════════════════════════════════

  Data-plane URL:  http://<host>:8788          ← your tools connect here
  Admin UI:        http://<host>:8788/admin    ← sign in here

  🔑 KEEP SAFE — lose these and you can be permanently locked out of your data:
     • CRYPTO_ENV_KEK_KEY  the KEK — encrypts every provider key you store
     • admin token         gw_admin_…          ← your Admin UI login
     • DB password         stored in the config
     • the gateway key (gw_sk_…) is minted later in the Admin UI / quickstart
```

* The **admin token** (`gw_admin_…`) is your Admin UI login — copy it now (it's shown once).
* The **KEK** (`CRYPTO_ENV_KEK_KEY`) is the most important secret to back up: it encrypts every provider key you'll add. Losing it means the stored provider keys can't be decrypted. **Copy it to your password manager.**

## Step 3: Sign in to the admin UI

1. Go to **`http://<host>:8788/admin`**, replacing **`<host>`** with the address of the machine running Governor.
2. Enter the admin token you recorded in Step 2.
3. Create a workspace and open it — your tenant (like a team/project).

<figure><img src="https://2860197046-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FYgNBTrPKG0DuVdAyDvSa%2Fuploads%2Ftqcb42qDbM8ZXET6R71a%2Fscrnli_6m8OC5RBA454Nu.png?alt=media&amp;token=5d264668-f648-4823-9051-32d8a72b2941" alt=""><figcaption></figcaption></figure>

The workspace opens on the **Dashboard**. The left sidebar contains the following pages: Dashboard, Documentation, Reports, Keys, Members, Accounts, Routes, Features, Limits, Prices, Test, and Audit.

<figure><img src="https://2860197046-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FYgNBTrPKG0DuVdAyDvSa%2Fuploads%2FCD72o1zTsd39AtgYDI37%2Fscrnli_zp2ysse2447CN9.png?alt=media&amp;token=4a3ae733-b564-4907-8bcc-6b4edafc7400" alt=""><figcaption></figcaption></figure>

Most configuration pages contain a form at the top and a table of existing entries below it. Complete Steps 4 to 9 in the admin UI.

{% hint style="info" %}
To perform the same configuration from a terminal instead, run `bito-gateway-ctl gwctl quickstart`.
{% endhint %}

## Step 4: Add a provider account

A provider account stores one LLM API key. Add an account for every LLM provider you use, such as Anthropic, OpenAI, Groq, Fireworks, OpenRouter, Together, Google, etc. Add more than one account for the same provider when you hold several keys, for example one per team or environment.

1. In the left sidebar, click **Accounts**.
2. In the **Add account** form, complete the following fields:

| Field    | Description                                                                                                                                     |
| -------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |
| provider | The provider you are connecting.                                                                                                                |
| name     | A name for this account, for example `prod` or `my-vllm`. The name appears in routes, reports, and prices.                                      |
| key      | Your provider API key. Governor stores it envelope encrypted under the key encryption key on your host.                                         |
| base url | The provider endpoint. The default is filled in for each provider. Override it for a proxy, a regional endpoint, or a self-hosted model server. |

3. Click **Add**.

The account appears in the **Accounts** table with the actions **edit**, **test**, **reveal**, and **remove**, and a toggle to enable or disable it.

4. Click **test** on the new account.

**test** checks the connection to the provider to verify the key and endpoint.

{% hint style="info" %}
To forward the credential supplied by the caller, leave the **key** field blank. This is passthrough mode.

Azure accounts can authenticate with an Entra ID service principal instead of an API key.
{% endhint %}

{% hint style="info" %}
**reveal** displays a stored provider key. Every use is recorded in the audit log.
{% endhint %}

#### Examples

<details>

<summary><strong>Expand to view examples</strong></summary>

Each example states a goal, then the values to enter in the **Add account** form.

#### Connect a provider

Add your organization's Anthropic key.

| provider  | name             | key                    | base url                  |
| --------- | ---------------- | ---------------------- | ------------------------- |
| anthropic | `prod-anthropic` | your Anthropic API key | leave the prefilled value |

Governor fills in **base url** when you select a provider. Change it only for a proxy, a regional endpoint, or a server you host yourself.

#### Separate production and staging spend

Bill staging traffic to a different key from production.

| provider  | name                | key                 | base url                  |
| --------- | ------------------- | ------------------- | ------------------------- |
| anthropic | `prod-anthropic`    | your production key | leave the prefilled value |
| anthropic | `staging-anthropic` | your staging key    | leave the prefilled value |

Add the provider twice under different names. Routes select an account by name, and reports show which account served each request, so spend separates cleanly.

#### Connect a model you host yourself

Send requests to a model running on your own inference server.

| provider | name      | key                                                | base url                      |
| -------- | --------- | -------------------------------------------------- | ----------------------------- |
| openai   | `my-vllm` | your server's key, or leave empty if it needs none | your inference server address |

Any server that implements the OpenAI API works here.

#### Forward the caller's own key

Apply routing and reporting to traffic without storing provider keys centrally.

| provider | name                 | key         | base url                  |
| -------- | -------------------- | ----------- | ------------------------- |
| openai   | `passthrough-openai` | leave empty | leave the prefilled value |

An account with no key runs in passthrough mode, where Governor forwards the credential the caller supplied.

</details>

## Step 5: Add a route

A route maps a model alias your coding tool requests to a provider account and an upstream model.

1. In the left sidebar, click **Routes**.
2. In the **Add route** form, complete the following fields:

<table data-search="false"><thead><tr><th align="center">Field</th><th>Description</th></tr></thead><tbody><tr><td align="center">alias</td><td>The model name your coding tool requests, for example <code>claude-opus-4-8</code> or <code>claude-*</code>. Enter <code>*</code> to match any model.</td></tr><tr><td align="center">provider</td><td>The provider that serves the request.</td></tr><tr><td align="center">account</td><td>The provider account used.</td></tr><tr><td align="center">model (or *)</td><td>The upstream model sent to the provider. Enter <code>*</code> to forward the model name the tool requested.</td></tr><tr><td align="center">api</td><td>The API dialect. Leave it on <strong>auto</strong> to follow the caller.</td></tr><tr><td align="center">priority</td><td>The failover tier. Lower numbers are used first. Default <code>0</code>.</td></tr><tr><td align="center">weight</td><td>The share of traffic within a tier. Default <code>1</code>.</td></tr><tr><td align="center">retries</td><td>Retry attempts for this target. Leave blank to use the gateway default.</td></tr><tr><td align="center">auto-route this alias</td><td>Adds the route, then opens auto-routing so you can say which models answer which kind of request. It starts as a dry run: every request is classified and recorded, and this route keeps answering until you turn routing on.<br><br>See <a href="/governor/auto-ai-model-routing.md">Auto AI model routing</a> for more details.</td></tr><tr><td align="center">backing group (internal)</td><td>Hides this route from your apps. They cannot ask for it by name and it does not appear in the model list, so only auto-routing can send traffic here. Use it for the models a router picks between, so nobody skips the router by asking for one directly.</td></tr></tbody></table>

3. Click **Add route**.

Routes are grouped by alias in the table below the form. Each group lists its targets with provider, model, account, priority, weight, retries, an enable toggle, and a health state. Priority, weight, and retries are editable inline.

Each group also carries a **Direct** and an **Auto** control. **Direct** sends every request for that alias to the targets you configured. **Auto** hands the alias to the auto-router, which picks a model for each request.

A target marked **cooling** is in a cooldown period after repeated failures. Governor sends its traffic to the next available target until it recovers.

#### Wildcard and exact aliases

An alias can be an exact model name such as `claude-opus-4-8`, or a wildcard such as `claude-sonnet-*` or `*`.

A route with `*` as both the alias and the upstream model passes every request through to the selected account unchanged.

Exact aliases take precedence over wildcards. A route for `claude-opus-4-8` is used instead of a `claude-*` route, and a `claude-*` route is used instead of `*`. This lets you run one catch-all route so that every model reaches a provider, then override selectively for the models you care about.

#### Failover and load balancing

Add the same alias again with a different provider account.

* Governor uses the lowest priority tier first.
* Within a tier, traffic is distributed by weight.
* Targets with the same priority receive requests in turn.
* If a provider returns errors, Governor sends the request to the next available target.

#### Auto-routing

An alias can answer from the targets you configured, or you can hand it to the auto-router, which classifies each request and picks a model for it. Select **auto-route this alias** when you add the route, or click **Auto** on the group in the routes table.

See [Auto AI model routing](/governor/auto-ai-model-routing.md) for more details.

#### Examples

<details>

<summary><strong>Expand to view examples</strong></summary>

Each example states a goal, then the values to enter in the **Add route** form.

#### Route every model to one provider

Send all traffic to a single account, without naming each model.

| alias | provider  | account          | model | priority |
| ----- | --------- | ---------------- | ----- | -------- |
| `*`   | anthropic | `prod-anthropic` | `*`   | 0        |

Type the asterisk in both fields. `*` in the alias matches any model, and `*` in the model field forwards the model name your tool sent. Add this route first, so that every request reaches a provider while you configure the rest.

#### Route a model family to one account

Send every version of Sonnet to an Azure account.

| alias             | provider | account      | model | priority |
| ----------------- | -------- | ------------ | ----- | -------- |
| `claude-sonnet-*` | azure    | `azure-prod` | `*`   | 0        |

The wildcard matches `claude-sonnet-5`, `claude-sonnet-4-5`, and later versions, so new releases need no new route. This alias is more specific than `*`, so it overrides the catch-all.

#### Fail over to a second provider

Serve Opus 5 from Anthropic, and switch to Azure while Anthropic is unavailable.

| alias           | provider  | account          | model           | priority |
| --------------- | --------- | ---------------- | --------------- | -------- |
| `claude-opus-5` | anthropic | `prod-anthropic` | `claude-opus-5` | 0        |
| `claude-opus-5` | azure     | `azure-prod`     | `claude-opus-5` | 1        |

Add the alias once per account. Priority `0` takes all traffic. When those requests fail, Governor moves them to priority `1` until the primary account recovers.

#### Split traffic across two accounts

Spread Opus 5 traffic so that neither account reaches its rate limit.

| alias           | provider  | account          | model           | priority | weight |
| --------------- | --------- | ---------------- | --------------- | -------- | ------ |
| `claude-opus-5` | anthropic | `prod-anthropic` | `claude-opus-5` | 0        | 3      |
| `claude-opus-5` | azure     | `azure-prod`     | `claude-opus-5` | 0        | 1      |

Equal priorities put both targets in the same tier, and the weights send three requests to Anthropic for every one to Azure. Set both weights to `1` for an even split.

#### Serve a request with a different model

Serve Opus 5 requests with a lower-cost model, with no change on any developer machine.

| alias           | provider  | account          | model             | priority |
| --------------- | --------- | ---------------- | ----------------- | -------- |
| `claude-opus-5` | anthropic | `prod-anthropic` | `claude-sonnet-5` | 0        |

The alias is the model your tool requests. The model is what serves the request. Your tools continue to request Opus 5, and Governor serves those requests with Sonnet 5. Edit the route to reverse it.

</details>

{% hint style="info" %}
Before applying a model substitution across your workspace, run it against a representative set of your own tasks and compare results as well as cost.
{% endhint %}

## Step 6: Create a gateway key

A gateway key authenticates a coding tool to Governor. Provider keys stay in the Bito UI.

1. In the left sidebar, click **Keys**.
2. In the **Create key** form, enter a **name** that identifies the team or tool that will use the key.
3. Click **Create**.
4. Copy the key.

The key starts with `gw_sk_` and is displayed once. Governor stores keys hashed and cannot display them again.

The **Keys** table lists each key by ID, prefix, and name, with an enable toggle and a **revoke** action. Create one key per team or tool so that you can revoke one without affecting the others.

## Step 7: Enable features

Features are server-side capabilities that Governor runs inside a request. Two features are available, and both are optional.

The **Features** table lists each enabled feature with its alias, state, whether a token is stored, and its MCP URL, with **edit**, **test connection**, and **disable** actions. Enabled features also appear as badges against each alias on the **Routes** page.

#### AI Architect

AI Architect serves system context from a live knowledge graph of your engineering system, covering code, business context, and tribal knowledge. Governor applies it inside each request, so your coding tools receive that context as they work.

1. In the left sidebar, click **Features**.
2. In the **Configure feature** form, select `ai_architect` from the **feature** list.
3. Complete the following fields:

<table data-search="false"><thead><tr><th>Field</th><th>Description</th></tr></thead><tbody><tr><td>alias</td><td>Leave blank to enable the feature across the workspace, or enter a route alias such as <code>claude-*</code> to scope it to that alias.</td></tr><tr><td>MCP URL</td><td>Your AI Architect MCP endpoint. If left empty, Governor falls back to a built-in stub.</td></tr><tr><td>Steering text</td><td>Overrides the default instructions Governor sends with AI Architect. Leave blank to use the default.</td></tr><tr><td>Tool allowlist</td><td>Restricts which AI Architect tools the model may call, one tool name per line. Leave empty to allow all.</td></tr><tr><td>Max hops</td><td>The hop budget for a single AI Architect lookup. Default <code>16</code>. See <a href="#max-hops">Max hops</a> below.</td></tr><tr><td>Max hops per request</td><td>The combined hop budget across every AI Architect lookup in one request. Leave blank to use the gateway default.</td></tr><tr><td>Prompt-cache injection</td><td>Caches what Governor sends with AI Architect on Anthropic, so repeat hops bill at the cache-read rate. Leave on <strong>Use default</strong> to follow the gateway-wide setting.</td></tr><tr><td>Architect model (sub-agent)</td><td>Runs the AI Architect lookups on a cost-efficient sub-agent while the route model writes the answer, for example GPT-5.6 Luna or Gemini Flash Lite 3.5. Leave blank to run them on the request's own route model.<br><br>You must configure a sub-agent to capture most of the cost savings.</td></tr><tr><td>Run Architect in-loop (legacy)</td><td>Runs the AI Architect tools inline on the route model instead of the default sub-agent mode. Ignored when an Architect model is set.</td></tr><tr><td>Delegation pressure</td><td><p>How hard the assistant is pushed to consult AI Architect.</p><ul><li><strong>Balanced</strong> consults it when the assistant judges research worthwhile, which suits coding tools whose own file search competes for the same job.</li><li><strong>Aggressive</strong> tells the assistant to consult AI Architect first, before searching files itself, which suits chat-style products.</li></ul><p>Left unset it follows the <strong>Architect model</strong> setting above — Balanced with none pinned, Aggressive with one, because a pinned model makes research cheap enough to lean on; the unset option names whichever applies right now. Choosing a value explicitly overrides that link and holds even if the Architect model changes later.</p></td></tr><tr><td>Evidence format</td><td><p>What a finished research call hands back.</p><ul><li><strong>Prose</strong> returns a written answer with citations.</li><li><strong>Structured</strong> returns the same research as separate findings, each carrying a verbatim code excerpt with its file and line.</li></ul><p>Prose is recommended, and matched or beat Structured on accuracy in five test setups out of six while costing less in five out of six, because excerpts add to the size of every request. Choose Structured when you parse the findings yourself.</p><p></p><p>Default: Prose</p></td></tr><tr><td>Quality mode</td><td><p>Sets how deeply AI Architect researches a question.<br><br>Choose one of the following:</p><ul><li><strong>Use default:</strong> follows the gateway-wide setting.</li><li><strong>Normal:</strong> uses the standard prompts and hop budget. This is the default.</li><li><strong>High quality:</strong> researches deeper and more thoroughly, at roughly twice the AI Architect cost.</li></ul></td></tr><tr><td>Codebase context</td><td><p>Codebase context tells the assistant about your own codebase before it starts work.</p><p></p><p>This is the highest-impact setting on this form, and it is off by default.</p><p></p><ul><li><strong>Conventions</strong> sends a short summary of the repository the person is working in, covering how you handle errors and logging, how you name things, how you test, and your security and module boundaries. The summary is the same for every request, so it is inexpensive to send repeatedly, and in testing it roughly halved the cost of a coding session.</li><li><strong>Conventions + task research</strong> also reads the request, works out what kind of work it is, and looks that up before answering, covering where the change belongs, what it would break, which pattern to follow, and what is already underway.</li></ul><p></p><p>Default: Off</p></td></tr><tr><td>Include risk areas and in-flight work</td><td><p>Also tells the assistant about known risk areas, technical debt, and work currently underway in the part of the codebase the caller is working in. Carries no contributor names.</p><p></p><p>Available only when <strong>Codebase context</strong> is on.</p></td></tr><tr><td>MCP token</td><td>Your AI Architect access token. Leave blank to keep the token already stored.</td></tr></tbody></table>

**Task guidance**

Task guidance adds what only your indexed codebase can supply to three kinds of request. Each setting is off by default and applies only to the request type named.

| Field                    | What it adds                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Guide what plans cover   | <p>When someone asks the assistant to plan or scope a change, it also covers what else the change would affect across your repositories, the patterns your code already follows, the business rules that constrain it, and the parts of the code that are risky to touch.</p><p></p><p>This makes plans more complete rather than more accurate, so it ensures the important ground is covered without making individual details more likely to be right.</p><p></p><p>In testing, plans that addressed those areas rose from about half to roughly four in five, with no slowdown.</p><p></p><p>Applies to planning, design, and implementation requests. Questions and troubleshooting are unaffected.</p> |
| Guide what reviews cover | <p>When someone asks what the codebase requires of a change they are about to make, the assistant also establishes which other repositories and services call the code being changed and would have to move with it, the convention the affected code is meant to follow, and the invariants the change could break.</p><p></p><p>The cross-repository part is the point, because call sites and conventions inside one repository are already searchable while a consumer in another service is not.</p><p></p><p>Applies to review requests only.</p>                                                                                                                                                      |
| Guide what triage covers | <p>When someone brings a bug or an outage, the assistant first checks what its local view cannot show, including whether an incident or alert is already firing for the services involved, what recently deployed in that area, and any issue already recorded against it.</p><p></p><p>Tracing the code is something the assistant can already do from the repository, while knowing that an alert has been firing since this morning is not.</p><p></p><p>Applies to bug reports and troubleshooting only.</p>                                                                                                                                                                                             |

4. Click **Enable**.

Changes apply to the next request. Users take no action.

{% hint style="info" %}
Setting **Architect model (sub-agent)** to a cost-efficient alias moves the AI Architect lookups off your main model while the route model still writes the answer. This reduces the cost of a request that takes several hops.
{% endhint %}

#### Max hops

Some questions require several passes to answer. A question about how your repositories connect requires Governor to retrieve the repository list, then look up the dependencies of each repository. Each pass is a hop.

Two settings cap this work, and both apply at the same time.

| Setting              | Scope                                               |
| -------------------- | --------------------------------------------------- |
| Max hops             | One AI Architect lookup.                            |
| Max hops per request | Every AI Architect lookup in one request, combined. |

A single request can trigger more than one lookup, so **Max hops per request** is what stops a complex request from running up cost through repeated lookups that each stay within their own limit.

Governor stops as soon as it has an answer, so both values are ceilings rather than fixed costs.

| Max hops     | Use                                                                                               |
| ------------ | ------------------------------------------------------------------------------------------------- |
| 16           | Default. Suitable for most workspaces.                                                            |
| 30 or higher | Workspaces with several hundred repositories, or teams that ask broad cross-repository questions. |

Raise **Max hops** if answers come back incomplete. Leave **Max hops per request** blank to use the gateway default, and set it when you want a firm ceiling on how much AI Architect work a single request can do. Each hop consumes tokens.

#### Reasoning downgrade

`reasoning_downgrade` lowers the reasoning effort of a request by exactly one level. Reasoning tokens bill at the output rate, so a lower level reduces the cost of the request.

1. In the left sidebar, click **Features**.
2. In the **Configure feature** form, select `reasoning_downgrade` from the **feature** list.
3. Leave **alias** blank to apply the feature across the workspace, or enter a route alias such as `claude-*` to scope it to that alias.
4. Click **Enable**.

Apart from **alias**, the feature has no settings.

Governor leaves a request unchanged when it already uses the lowest or second-lowest reasoning level, or when it sends no reasoning at all.

## Step 8: Connect a coding tool

In the left sidebar, click **Documentation**. Your own base URL is displayed at the top of the page, and the links below it jump to the four sections on the page.

| Section        | Contents                                                                                     |
| -------------- | -------------------------------------------------------------------------------------------- |
| Get started    | Your base URL, a field for your gateway key, and how to authenticate.                        |
| Use it         | Ready-made curl, Python, and JavaScript requests, and **Try it** for sending a live request. |
| Connect a tool | Copy-paste setup for each supported coding tool.                                             |
| Reference      | Endpoints, your route aliases, and how usage and cost are reported.                          |

To connect a coding tool:

1. In the **Get started** section, paste a gateway key into the key field. Every example on the page fills in with your base URL and key. To generate a key here, click **+ Create test key**.
2. In the **Connect a tool** section, select your tool.

Setup is provided for Claude Code, Cursor, Cline (VS Code), Continue (VS Code / JetBrains), Aider, Codex CLI, GitHub Copilot CLI, Windsurf, Zed, and any OpenAI-compatible or Anthropic-compatible tool.

To print the same setup from a terminal, run `bito-gateway-ctl gwctl connect` with your tool name. The command outputs configuration values and changes no state.

```bash
bito-gateway-ctl gwctl connect claude
# Supported values: claude | cursor | cline | continue | aider | codex | copilot | windsurf | zed | openai | anthropic
```

The output contains your data-plane URL and the required environment variables, with a `<YOUR_GATEWAY_KEY>` placeholder. Replace it with the key you created in Step 6.

For Claude Code, set two environment variables and start the tool:

```bash
export ANTHROPIC_BASE_URL="https://HOST:8788"
export ANTHROPIC_AUTH_TOKEN="gw_sk_..."
claude
```

Replace `HOST` with the hostname or address of the machine running Governor. To configure a team, distribute these two variables using your existing developer environment tooling. If your traffic already passes through a central gateway, set them there instead.

Governor accepts requests on three endpoints:

| Path                   | Dialect                 | Used by                             |
| ---------------------- | ----------------------- | ----------------------------------- |
| `/v1/messages`         | Anthropic Messages      | Claude Code, Anthropic SDKs         |
| `/v1/chat/completions` | OpenAI Chat Completions | Codex, most OpenAI-compatible tools |
| `/v1/responses`        | OpenAI Responses        | OpenAI Responses API clients        |

Authenticate with `Authorization: Bearer <GATEWAY_KEY>` or `X-Api-Key: <GATEWAY_KEY>`. The same gateway key works for all three dialects.

Every response carries an `x-bito-routed-model` header naming the model that answered. Read that header rather than the `model` field inside the response body, because some providers return the model that was requested and others return the model that answered. The difference matters once auto-routing is on, since the model that answers may not be the one your tool asked for.

{% hint style="info" %}
Claude Code reports that its connectors are disabled when it runs through Governor, because the session authenticates against Governor rather than a Claude account. This message is expected. AI Architect continues to work, because Governor serves it from the server side.
{% endhint %}

## Operate the Bito Governor

Day-to-day, use the **`bito-gateway-ctl`** command (installed on your PATH — Windows too, via a `.cmd` shim). It works from any directory:

```bash
bito-gateway-ctl status         # is it up?
bito-gateway-ctl health         # /healthz + a deep self-check (gwctl doctor)
bito-gateway-ctl logs           # follow the gateway logs (Ctrl-C to stop)
bito-gateway-ctl restart
bito-gateway-ctl stop           # stop, keep your data
bito-gateway-ctl start          # start again (also applies .env changes)
bito-gateway-ctl update         # pull a newer image and restart
```

**Run `gwctl` from the command line** (it lives inside the gateway container — this wrapper forwards to it):

```bash
bito-gateway-ctl gwctl quickstart              # interactive: workspace + key + account + route (the CLI version of §4)
bito-gateway-ctl gwctl connect claude          # print a tool's setup
bito-gateway-ctl gwctl admin create --role global   # add another admin
bito-gateway-ctl gwctl doctor                  # validate DB/schema/KEK
```

**Uninstall:**

```bash
bito-gateway-ctl uninstall           # remove the containers, KEEP your data (the DB volume + secrets)
bito-gateway-ctl uninstall --purge   # ALSO delete the DB volume + secrets (the KEK) — irreversible
```

{% hint style="info" %}
On the **native/systemd** install, use `sudo systemctl {start|stop|restart|status} bito-gateway` + `journalctl -u bito-gateway -f`, and `gwctl` is already on your PATH directly.
{% endhint %}

### Verify the configuration

1. **Check route selection.** In the left sidebar, click **Test**. Enter a model alias and click **Test**. Governor returns the route, provider, account, and lane it would use. No request is sent, so this costs nothing.
2. **Check auto-routing, if you set it up.** In the left sidebar, click **Routes**. On a routed alias, click **Classify a prompt**, enter a sample request, and confirm it lands in the tier you expect. No request is sent, so this costs nothing.
3. **Send a request.** In the left sidebar, click **Documentation**. In the **Use it** section, go to **Try it**. Paste a gateway key, select a model, type a prompt in the message box, and click **Run**. This sends a real, billable request to your provider.
4. **Query your own system.** From your coding tool, ask a question that requires knowledge of your repositories. A response naming your own services confirms AI Architect is active.
5. **Check the report.** In the left sidebar, click **Reports** and confirm the requests appear.

{% hint style="info" %}
For a wildcard alias such as `claude-*`, enter a concrete model that matches it, for example `claude-opus-4-8`.
{% endhint %}

### Set prices

Token prices produce the cost figures on the **Dashboard** and in **Reports**. Prices are expressed in `$/Mtok`, meaning US dollars per million tokens.

Global defaults are maintained by your gateway operator. Set a price here to override the default for your workspace, or to record a negotiated rate for one account. Governor resolves prices in this order: account, then workspace, then global.

1. In the left sidebar, click **Prices**.
2. In the **Set price** form, select the **vendor**.
3. Select an **account** to apply the price to that account only, or leave it on **all accounts** to apply the vendor rate across your workspace.
4. Enter the **model** name.
5. Enter **in**, **out**, and optionally **cache-read** and **cache-write** prices, all in `$/Mtok`.
6. Click **Set**.

The **Prices** table lists each price with its scope, vendor, model, rates, and source.

A model with no price reports token counts and a cost of zero. The **Dashboard** shows the number of unpriced requests per model in the **UNPRICED** column. Auto-routing also depends on these prices. A model with no price cannot be selected by a router, and both the routing direction and the price ceiling are evaluated against them.

### Set limits

Limits are applied per gateway key.

1. In the left sidebar, click **Limits**.
2. In the **Set limits** form, select a **key**.
3. Complete any of the following fields:

| Field            | Description                                   |
| ---------------- | --------------------------------------------- |
| rpm              | Maximum requests per minute.                  |
| tpm              | Maximum tokens per minute.                    |
| budget ($/month) | Monthly spend cap. Requires prices to be set. |
| max\_concurrency | Maximum concurrent requests.                  |

4. Click **Set**.

Leave a field blank to leave it unchanged. Enter `0` to remove a cap, which makes that limit unlimited.

{% hint style="info" %}
Setting a limit to `0` removes the cap rather than blocking the key. To stop a key entirely, disable it on the **Keys** page.
{% endhint %}

### Monitor usage and cost

#### Dashboard

The **Dashboard** shows spend and request volume for the last 30 days, with totals for requests, input tokens, output tokens, and cost. Use **Group by** to break the numbers down by model, alias, key, provider, account, detail, or feature.

The table below the chart lists requests, token counts, cost, and unpriced request count per model.

#### Reports

**Reports** is the raw event log, with one row per request, updated in near real time. Filter the log and export it with **Download CSV**.

Token counts are split into four buckets:

| Bucket   | Description                                                             |
| -------- | ----------------------------------------------------------------------- |
| in       | Fresh prompt tokens, excluding anything served from cache.              |
| cached   | Prompt tokens read from cache, billed at a lower rate.                  |
| cache\_w | Tokens written to the cache. On Anthropic this carries a small premium. |
| out      | All generated tokens, including reasoning tokens.                       |

The full prompt your tool sent is `in + cached + cache_w`. The buckets do not overlap, so nothing is counted twice.

Features make hidden calls inside a request and are reported separately. The base columns show the answer the caller received, and the feature columns show the feature's own usage. A request's total is base plus feature. Expand a row to see the split.

#### Measure the effect of AI Architect

Governor does not currently report savings against a baseline. To measure the effect:

1. Select a set of tasks your team runs regularly.
2. Disable AI Architect, run the tasks, and record cost per task from **Reports**.
3. Enable AI Architect and run the same tasks with the same tool and model.
4. Compare cost per task and confirm the tasks still complete correctly.

### Add members

1. In the left sidebar, click **Members**.
2. In the **Add team member** form, enter a **name** and select a **role**.
3. Click **Create**.

Each member signs in to the admin UI with the token generated for them.

| Role            | Permissions                                                                                   |
| --------------- | --------------------------------------------------------------------------------------------- |
| Member          | Manages their own gateway keys and views their own usage.                                     |
| Workspace admin | Full configuration access to the workspace, including routes, accounts, features, and limits. |

You can disable a member from the **Team** table.

### View the audit log

**Audit** records every configuration change and every secret reveal in your workspace, with the admin who performed it and a timestamp. The log is read only.

Give each person their own member token, so that the log identifies who made each change.

***

### Back up

Two things, together, are your whole Bito Governor:

* **The database** — `mysqldump bito_gateway` on a schedule (it holds every workspace, key, account, route, and your encrypted provider secrets).
* **The KEK** (`CRYPTO_ENV_KEK_KEY`, in `~/.bito-gateway/.env`) — in your secrets manager. **A database dump alone can't be decrypted without it.**

***

## Troubleshooting

<table data-search="false"><thead><tr><th>Symptom</th><th>Cause</th><th>Resolution</th></tr></thead><tbody><tr><td><code>401</code> from Governor with a gateway key</td><td>The gateway key is invalid or revoked.</td><td>Create a new key on the <strong>Keys</strong> page and update the tool configuration.</td></tr><tr><td><code>404</code> for a model</td><td>No route matches the requested alias.</td><td>Add a route for the exact model name, or add a <code>*</code> route.</td></tr><tr><td>Requests reach an unexpected model</td><td>An exact alias takes precedence over the <code>*</code> route.</td><td>Check the <strong>Routes</strong> page. Exact aliases override wildcards.</td></tr><tr><td>A route shows <strong>cooling</strong></td><td>The target failed repeatedly and is in a cooldown.</td><td>Check the provider account with <strong>test</strong> on the <strong>Accounts</strong> page. Traffic uses the next target until it recovers.</td></tr><tr><td>Cost column shows zero, or UNPRICED is high</td><td>The model has no price set.</td><td>Add prices for that model on the <strong>Prices</strong> page.</td></tr><tr><td>Budget limit has no effect</td><td>The model has no price set, or the limit is <code>0</code>.</td><td>Set prices, then set a positive budget. <code>0</code> means unlimited.</td></tr><tr><td>A key still works after setting limits to <code>0</code></td><td><code>0</code> removes the cap rather than blocking the key.</td><td>Disable the key on the <strong>Keys</strong> page.</td></tr><tr><td>AI Architect responses lack system context</td><td>The feature is disabled, scoped to a different alias, or the MCP URL is empty.</td><td>Open the <strong>Features</strong> page and check the alias, MCP URL, and token. An empty MCP URL falls back to a built-in stub.</td></tr><tr><td>Broad questions return incomplete answers</td><td>Governor reached the max hops budget.</td><td>Increase <strong>Max hops</strong>, and check <strong>Max hops per request</strong> if the request makes several AI Architect lookups.</td></tr><tr><td>Repeated <code>429</code> or <code>5xx</code></td><td>One provider account is rate limited or unavailable.</td><td>Add a second target on the same alias to enable failover.</td></tr><tr><td>Requests reach a model you did not configure for that alias</td><td>Auto-routing is on for the alias.</td><td>Open <strong>Routes</strong> and check the tier assignments, or click <strong>Direct</strong> to stop routing that alias.</td></tr><tr><td>A check that compares the requested model to the returned model fails</td><td>The <code>model</code> field in the response names the model that answered on some providers.</td><td>Read the <code>x-bito-routed-model</code> response header instead, which always names the model that answered.</td></tr><tr><td>Auto-routing is on but nothing is routed</td><td>The router is in shadow, or every tier is unusable.</td><td>Check the mode on <strong>Routes</strong>. A tier marked <strong>will not route</strong> has a model with no price, no deployment on the account, or one blocked by your routing direction.</td></tr><tr><td>Every request is classified simple</td><td>The LLM classifier is timing out, so the built-in scorer is deciding instead.</td><td>Raise the classifier <strong>timeout</strong>, or point <strong>classifier model</strong> at a faster route alias.</td></tr></tbody></table>

## What's next

* The Admin UI **Documentation** tab — connection snippets for every tool, filled with your live URL/key, plus the API reference.
* `bito-gateway-ctl gwctl <command> -h` — help for any CLI command.
* [Bito Governor overview](/governor/overview.md)
* [Set up Bito Governor (Bito-hosted)](/governor/set-up-bito-governor-bito-hosted.md)
* [Set up Bito Governor (self-hosted with Kubernetes)](/governor/set-up-bito-governor-self-hosted-with-kubernetes.md)
* Contact <support@bito.ai> for configuration assistance.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.bito.ai/governor/set-up-bito-governor-self-hosted-with-docker.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
