Set up Bito Governor (Bito-hosted)
Get your team running on Bito-hosted Governor. Connect your LLM providers, set up routing and gateway keys, enable AI Architect, and point your coding tools at it.
Bito Governor runs either on Bito's infrastructure or inside your own network. This page covers the Bito-hosted deployment. To run Governor inside your own network, see Set up Bito Governor (self-hosted).
With Bito-hosted Governor, Bito operates the infrastructure and there is nothing for you to install. You configure your workspace in the Bito Governor admin UI at https://gateway.bito.ai/admin/, then point your AI coding tools at it.
Setup has four parts:
Provider account
Stores an API key for an LLM provider. Add an account for every provider you use.
Route
Which provider and model serve each request.
Maps a model alias your tools request to a provider account and an upstream model.
Gateway key
Authenticates a coding tool to Governor.
Feature
A server-side capability that runs inside a request, such as Bito's AI Architect. Optional.
Prerequisites
You need the following:
Admin Token. Contact support@bito.ai to set up Governor for the first time. Bito creates your account and sends you an Admin Token.
An API key for each LLM provider you use, such as Anthropic, OpenAI, Groq, or Google. Governor sends requests to your own provider accounts.
The model names your coding tools request, such as
claude-opus-4-8. Model names are case sensitive.AI Architect MCP URL and access token, if you enable AI Architect. Both are available in your Bito account.
Sign in
Enter your Admin Token.
The workspace opens on the Dashboard. The left sidebar contains the following pages: Dashboard, Documentation, Reports, Keys, Members, Accounts, Routes, Features, Limits, Prices, Test, and Audit.
Your Admin Token grants access to your own workspace only.
Most configuration pages contain a form at the top and a table of existing entries below it.
Step 1: Add a provider account
A provider account stores one API key. Add an account for every provider you use, such as Anthropic, OpenAI, Groq, Fireworks, OpenRouter, Together, or Google. Add more than one account for the same provider when you hold several keys, for example one per team or environment.
In the left sidebar, click Accounts.
In the Add account form, complete the following fields:
provider
The provider you are connecting.
name
A name for this account, for example prod or my-vllm. The name appears in routes, reports, and prices.
key
Your provider API key. Governor stores it envelope encrypted.
base url
The provider endpoint. The default is filled in for each provider. Override it for a proxy, a regional endpoint, or a self-hosted model server.
Click Add.
The account appears in the Accounts table with the actions edit, test, reveal, and remove.
Click test on the new account. test checks the connection to the provider to verify the key and endpoint.
To forward the credential supplied by the caller, leave the key field blank. This is passthrough mode.
Azure accounts can authenticate with an Entra ID service principal instead of an API key.
reveal displays a stored provider key. Every use is recorded in the audit log.
Examples
Expand to view examples
Each example states a goal, then the values to enter in the Add account form.
Connect a provider
Add your organisation's Anthropic key.
anthropic
prod-anthropic
your Anthropic API key
leave the prefilled value
Governor fills in base url when you select a provider. Change it only for a proxy, a regional endpoint, or a server you host yourself.
Separate production and staging spend
Bill staging traffic to a different key from production.
anthropic
prod-anthropic
your production key
leave the prefilled value
anthropic
staging-anthropic
your staging key
leave the prefilled value
Add the provider twice under different names. Routes select an account by name, and reports show which account served each request, so spend separates cleanly.
Connect a model you host yourself
Send requests to a model running on your own inference server.
openai
my-vllm
your server's key, or leave empty if it needs none
your inference server address
Any server that implements the OpenAI API works here.
Forward the caller's own key
Apply routing and reporting to traffic without storing provider keys centrally.
openai
passthrough-openai
leave empty
leave the prefilled value
An account with no key runs in passthrough mode, where Governor forwards the credential the caller supplied.
Step 2: Add a route
A route maps a model alias your coding tool requests to a provider account and an upstream model.
In the left sidebar, click Routes.
In the Add route form, complete the following fields:
alias
The model name your coding tool requests, for example claude-opus-4-8 or claude-*. Enter * to match any model.
provider
The provider that serves the request.
account
The provider account used.
model (or *)
The upstream model sent to the provider. Enter * to forward the model name the tool requested.
api
The API dialect. Leave it on auto to follow the caller.
priority
The failover tier. Lower numbers are used first. Default 0.
weight
The share of traffic within a tier. Default 1.
retries
Retry attempts for this target. Leave blank to use the gateway default.
Click Add route.
Routes are grouped by alias in the table below the form. Each group lists its targets with provider, model, account, priority, weight, retries, an enable toggle, and a health state. Priority, weight, and retries are editable inline.
A target marked cooling is in a cooldown period after repeated failures. Governor sends its traffic to the next available target until it recovers.
Wildcard and exact aliases
A route with * as both the alias and the upstream model passes every request through to the selected account unchanged.
Exact aliases take precedence over wildcards. A route for claude-opus-4-8 is used instead of the * route, and all other models continue to pass through.
Failover and load balancing
Add the same alias again with a different provider account.
Governor uses the lowest priority tier first.
Within a tier, traffic is distributed by weight.
Targets with the same priority receive requests in turn.
If a provider returns errors, Governor sends the request to the next available target.
Examples
Expand to view examples
Each example states a goal, then the values to enter in the Add route form.
Route every model to one provider
Send all traffic to a single account, without naming each model.
*
anthropic
prod-anthropic
*
0
Type the asterisk in both fields. * in the alias matches any model, and * in the model field forwards the model name your tool sent. Add this route first, so that every request reaches a provider while you configure the rest.
Route a model family to one account
Send every version of Sonnet to an Azure account.
claude-sonnet-*
azure
azure-prod
*
0
The wildcard matches claude-sonnet-5, claude-sonnet-4-5, and later versions, so new releases need no new route. This alias is more specific than *, so it overrides the catch-all.
Fail over to a second provider
Serve Opus 5 from Anthropic, and switch to Azure while Anthropic is unavailable.
claude-opus-5
anthropic
prod-anthropic
claude-opus-5
0
claude-opus-5
azure
azure-prod
claude-opus-5
1
Add the alias once per account. Priority 0 takes all traffic. When those requests fail, Governor moves them to priority 1 until the primary account recovers.
Split traffic across two accounts
Spread Opus 5 traffic so that neither account reaches its rate limit.
claude-opus-5
anthropic
prod-anthropic
claude-opus-5
0
3
claude-opus-5
azure
azure-prod
claude-opus-5
0
1
Equal priorities put both targets in the same tier, and the weights send three requests to Anthropic for every one to Azure. Set both weights to 1 for an even split.
Serve a request with a different model
Serve Opus 5 requests with a lower-cost model, with no change on any developer machine.
claude-opus-5
anthropic
prod-anthropic
claude-sonnet-5
0
The alias is the model your tool requests. The model is what serves the request. Your tools continue to request Opus 5, and Governor serves those requests with Sonnet 5. Edit the route to reverse it.
Step 3: Create a gateway key
A gateway key authenticates a coding tool to Governor. Provider keys stay in the Bito UI.
In the left sidebar, click Keys.
In the Create key form, enter a name that identifies the team or tool that will use the key.
Click Create.
Copy the key.
The key starts with gw_sk_ and is displayed once. Governor stores keys hashed and cannot display them again.
The Keys table lists each key by ID, prefix, name, enabled status, and a revoke action.
Create one key per team or tool so that you can revoke one without affecting the others.
Step 4: Enable features
Features are server-side capabilities that Governor runs inside a request. Two features are available, and both are optional.
The Features table lists each enabled feature with its alias, state, whether a token is stored, and its MCP URL, with edit and disable actions. Enabled features also appear as badges against each alias on the Routes page.
AI Architect
AI Architect serves system context from a live knowledge graph of your engineering system, covering code, business context, and tribal knowledge. Governor applies it inside each request, so your coding tools receive that context as they work.
In the left sidebar, click Features.
In the Configure feature form, select
ai_architectfrom the feature list.Complete the following fields:
alias
Leave blank to enable the feature across the workspace, or enter a route alias such as claude-* to scope it to that alias.
MCP URL
Your AI Architect MCP endpoint. If left empty, Governor falls back to a built-in stub.
Steering text
Overrides the default instructions Governor sends with AI Architect. Leave blank to use the default.
Tool allowlist
Restricts which AI Architect tools the model may call, one tool name per line. Leave empty to allow all.
Max hops
The server-side hidden-loop hop budget. Default 16. See Max hops below for more details.
Max hops per request
The combined hop budget across every AI Architect lookup in one request. Leave blank to use the gateway default.
Prompt-cache injection
Caches what Governor sends with AI Architect on Anthropic, so repeat hops bill at the cache-read rate. Leave on Use default to follow the gateway-wide setting.
Architect model (sub-agent)
Runs the AI Architect lookups on a cheaper route alias while the route model writes the answer, for example gemini-3.1-flash-lite or gpt-5.6-luna. Leave blank to run them on the request's own route model.
Run Architect in-loop (legacy)
Runs the AI Architect tools inline on the route model instead of the default sub-agent mode. Ignored when an Architect model is set.
Quality mode
Sets how deeply AI Architect researches a question. Choose one of the following:
Use default: follows the gateway-wide setting.
Normal: uses the standard prompts and hop budget. This is the default.
High quality: researches deeper and more thoroughly, at roughly twice the AI Architect cost.
MCP token
Your AI Architect access token. Leave blank to keep the token already stored.
Click Enable.
Changes apply to the next request. Users take no action.
Setting Architect model (sub-agent) to a cheaper alias moves the AI Architect lookups off your main model while the route model still writes the answer. This reduces the cost of a request that takes several hops.
Max hops
Some questions require several passes to answer. A question about how your repositories connect requires Governor to retrieve the repository list, then look up the dependencies of each repository. Each pass is a hop.
Two settings cap this work, and both apply at the same time.
Max hops
One AI Architect lookup.
Max hops per request
Every AI Architect lookup in one request, combined.
A single request can trigger more than one lookup, so Max hops per request is what stops a complex request from running up cost through repeated lookups that each stay within their own limit.
Governor stops as soon as it has an answer, so both values are ceilings rather than fixed costs.
16
Default. Suitable for most workspaces.
30 or higher
Workspaces with several hundred repositories, or teams that ask broad cross-repository questions.
Raise Max hops if answers come back incomplete. Leave Max hops per request blank to use the gateway default, and set it when you want a firm ceiling on how much AI Architect work a single request can do. Each hop consumes tokens.
Reasoning downgrade
reasoning_downgrade lowers the reasoning effort of a request by exactly one level. Reasoning tokens bill at the output rate, so a lower level reduces the cost of the request.
In the left sidebar, click Features.
In the Configure feature form, select
reasoning_downgradefrom the feature list.Leave alias blank to apply the feature across the workspace, or enter a route alias such as
claude-*to scope it to that alias.Click Enable.
Apart from alias, the reasoning_downgrade feature has no settings.
Governor leaves a request unchanged when it already uses the lowest or second-lowest reasoning level, or when it sends no reasoning at all.
Step 5: Connect a coding tool
In the left sidebar, click Documentation. Your base URL is displayed at the top of the page, and the links below it jump to the four sections on the page.
Get started
Your base URL, a field for your gateway key, and how to authenticate.
Use it
Ready-made curl, Python, and JavaScript requests, and Try it for sending a live request.
Connect a tool
Copy-paste setup for each supported coding tool.
Reference
Endpoints, your route aliases, and how usage and cost are reported.
To connect a coding tool:
In the Get started section, paste a gateway key into the key field. Every example on the page fills in with your base URL and key. To generate a key here, click + Create test key.
In the Connect a tool section, select your tool.
Setup is provided for Claude Code, Cursor, Cline (VS Code), Continue (VS Code / JetBrains), Aider, Codex CLI, GitHub Copilot CLI, Windsurf, Zed, and any OpenAI-compatible or Anthropic-compatible tool.
For Claude Code, set two environment variables and start the tool:
To configure a team, distribute these two variables using your existing developer environment tooling. If your traffic already passes through a central gateway, set them there instead.
Governor accepts requests on three endpoints:
/v1/messages
Anthropic Messages
Claude Code, Anthropic SDKs
/v1/chat/completions
OpenAI Chat Completions
Codex, most OpenAI-compatible tools
/v1/responses
OpenAI Responses
OpenAI Responses API clients
Authenticate with Authorization: Bearer <GATEWAY_KEY> or X-Api-Key: <GATEWAY_KEY>. The same gateway key works for all three dialects.
Claude Code reports that its connectors are disabled when it runs through Governor, because the session authenticates against Governor rather than a Claude account. This message is expected. AI Architect continues to work, because Governor serves it from the server side.
Operate the Bito Governor
Verify the configuration
Check route selection. In the left sidebar, click Test. Enter a model alias and click Test. Governor returns the route, provider, account, and lane it would use. No request is sent, so this costs nothing.
Send a request. In the left sidebar, click Documentation. In the Use it section, go to Try it. Paste a gateway key, select a model, type a prompt in the message box, and click Run. This sends a real, billable request to your provider.
Query your own system. From your coding tool, ask a question that requires knowledge of your repositories. A response naming your own services confirms AI Architect is active.
Check the report. In the left sidebar, click Reports and confirm the requests appear.
For a wildcard alias such as claude-*, enter a concrete model that matches it, for example claude-opus-4-8.
Set prices
Token prices produce the cost figures on the Dashboard and in Reports. Prices are expressed in $/Mtok, meaning US dollars per million tokens.
Global defaults are maintained by your gateway operator. Set a price here to override the default for your workspace, or to record a negotiated rate for one account. Governor resolves prices in this order: account, then workspace, then global.
In the left sidebar, click Prices.
In the Set price form, select the vendor.
Select an account to apply the price to that account only, or leave it on all accounts to apply the vendor rate across your workspace.
Enter the model name.
Enter in, out, and optionally cache-read and cache-write prices, all in
$/Mtok.Click Set.
The Prices table lists each price with its scope, vendor, model, rates, and source.
A model with no price reports token counts and a cost of zero. The Dashboard shows the number of unpriced requests per model in the UNPRICED column.
Set limits
Limits are applied per gateway key.
In the left sidebar, click Limits.
In the Set limits form, select a key.
Complete any of the following fields:
rpm
Maximum requests per minute.
tpm
Maximum tokens per minute.
budget ($/month)
Monthly spend cap. Requires prices to be set.
max_concurrency
Maximum concurrent requests.
Click Set.
Leave a field blank to leave it unchanged. Enter 0 to remove a cap, which makes that limit unlimited.
Setting a limit to 0 removes the cap rather than blocking the key. To stop a key entirely, disable it on the Keys page.
Monitor usage and cost
Dashboard
The Dashboard shows spend and request volume for the last 30 days, with totals for requests, input tokens, output tokens, and cost. Use Group by to break the numbers down by model, alias, key, provider, account, detail, or feature.
The table below the chart lists requests, token counts, cost, and unpriced request count per model.
Reports
Reports is the raw event log, with one row per request, updated in near real time. Filter the log and export it with Download CSV.
Token counts are split into four buckets:
in
Fresh prompt tokens, excluding anything served from cache.
cached
Prompt tokens read from cache, billed at a lower rate.
cache_w
Tokens written to the cache. On Anthropic this carries a small premium.
out
All generated tokens, including reasoning tokens.
The full prompt your tool sent is in + cached + cache_w. The buckets do not overlap, so nothing is counted twice.
Features make hidden calls inside a request and are reported separately. The base columns show the answer the caller received, and the feature columns show the feature's own usage. A request's total is base plus feature. Expand a row to see the split.
Measure the effect of AI Architect
Governor does not currently report savings against a baseline. To measure the effect:
Select a set of tasks your team runs regularly.
Disable AI Architect, run the tasks, and record cost per task from Reports.
Enable AI Architect and run the same tasks with the same tool and model.
Compare cost per task and confirm the tasks still complete correctly.
Add members
In the left sidebar, click Members.
In the Add team member form, enter a name and select a role.
Click Create.
Each member signs in to the Bito UI with the token generated for them.
Member
Manages their own gateway keys and views their own usage.
Workspace admin
Full configuration access to the workspace, including routes, accounts, features, and limits.
You can disable a member from the Team table. To change or disable a workspace admin, contact support@bito.ai.
View the audit log
Audit records every configuration change and every secret reveal in your workspace, with the admin who performed it and a timestamp. The log is read only.
Give each person their own member token, so that the log identifies who made each change.
Troubleshooting
401 from Governor
The gateway key is invalid or revoked.
Create a new key on Keys and update the tool configuration.
404 for a model
No route matches the requested alias.
Add a route for the exact model name, or add a * route.
Requests reach an unexpected model
An exact alias takes precedence over the * route.
Open Routes. Exact aliases override wildcards.
A route shows cooling
The target failed repeatedly and is in a cooldown.
Check the provider account with test on the Accounts page. Traffic uses the next target until it recovers.
Cost column shows zero, or UNPRICED is high
The model has no price set.
Add prices for that model on Prices.
Budget limit has no effect
The model has no price set, or the limit is 0.
Set prices, then set a positive budget. 0 means unlimited.
A key still works after setting limits to 0
0 removes the cap rather than blocking the key.
Disable the key on Keys.
AI Architect responses lack system context
The feature is disabled, scoped to a different alias, or the MCP URL is empty.
Open Features and check the alias, MCP URL, and token. An empty MCP URL falls back to a built-in stub.
Broad questions return incomplete answers
Governor reached the max hops budget.
Increase Max hops and repeat the question.
Repeated 429 or 5xx
One provider account is rate limited or unavailable.
Add a second target on the same alias to enable failover.
What's next
Contact support@bito.ai for configuration assistance.
Last updated

