> ## Documentation Index
> Fetch the complete documentation index at: https://xum.cdr.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Providers

> Configure API keys and settings for AI providers

Xum supports multiple AI providers. The easiest way to configure them is through **Settings → Providers** (`Cmd+,` / `Ctrl+,`).

## Quick Setup

1. Open Settings (`Cmd+,` / `Ctrl+,`)
2. Navigate to **Providers**
3. Expand any provider and enter your API key
4. Start using models from that provider

Most providers only need an API key. The UI handles validation and shows which providers are configured.

## Supported Providers

| Provider | Models | Get API Key |
| - | - | - |
| **Anthropic** | Claude Opus, Sonnet, Haiku | [console.anthropic.com](https://console.anthropic.com/) |
| **OpenAI** | GPT-6 | [platform.openai.com](https://platform.openai.com/) |
| **Google** | Gemini Pro, Flash | [aistudio.google.com](https://aistudio.google.com/) |
| **xAI** | Grok | [console.x.ai](https://console.x.ai/) |
| **DeepSeek** | DeepSeek Chat, Reasoner | [platform.deepseek.com](https://platform.deepseek.com/) |
| **Moonshot AI** | Kimi K3 | [platform.moonshot.ai](https://platform.moonshot.ai/) |
| **Z.ai** | GLM 5.3 Flash | [z.ai](https://z.ai/model-api) |
| **OpenRouter** | 300+ models | [openrouter.ai](https://openrouter.ai/) |
| **Ollama** | Local models | [ollama.com](https://ollama.com/) (no key needed) |
| **Bedrock** | Claude via AWS | AWS Console |
| **GitHub Copilot** | GPT-4o, Claude Sonnet, etc. | [GitHub Copilot](https://github.com/features/copilot) |
| **Coder** | Models via AI Gateway | Your Coder deployment (Login with Coder, no key needed) |

For catalog suggestions and manual model IDs, see [Add custom models](/config/models#add-custom-models).

## Environment Variables

Providers also read from environment variables as fallback:

| Provider | Environment Variable |
| - | - |
| Anthropic | `ANTHROPIC_API_KEY` or `ANTHROPIC_AUTH_TOKEN` |
| OpenAI | `OPENAI_API_KEY` |
| Google | `GOOGLE_GENERATIVE_AI_API_KEY` or `GOOGLE_API_KEY` |
| xAI | `XAI_API_KEY` |
| OpenRouter | `OPENROUTER_API_KEY` |
| DeepSeek | `DEEPSEEK_API_KEY` |
| Moonshot AI | `MOONSHOT_API_KEY` |
| Z.ai | `ZAI_API_KEY` |
| github-copilot | `GITHUB_COPILOT_TOKEN` |
| Bedrock | `AWS_REGION` (credentials via AWS SDK chain) |

<details>
  <summary>Additional environment variables</summary>

  | Provider | Variable | Purpose |
  | - | - | - |
  | Anthropic | `ANTHROPIC_BASE_URL` | Custom API endpoint |
  | OpenAI | `OPENAI_BASE_URL` | Custom API endpoint |
  | OpenAI | `OPENAI_ORG_ID` | Organization ID |
  | Google | `GOOGLE_BASE_URL` | Custom API endpoint |
  | xAI | `XAI_BASE_URL` | Custom API endpoint |
  | Azure OpenAI | `AZURE_OPENAI_API_KEY` | API key |
  | Azure OpenAI | `AZURE_OPENAI_ENDPOINT` | Endpoint URL |
  | Azure OpenAI | `AZURE_OPENAI_DEPLOYMENT` | Deployment name |
  | Azure OpenAI | `AZURE_OPENAI_API_VERSION` | API version |

  Azure OpenAI env vars configure the OpenAI provider with Azure backend.
</details>

## AI Calls from Bash Commands

Scripts that the agent runs in the bash tool (test harnesses, `make bug-bash`, SDK scripts) can call Anthropic and OpenAI directly. Turn on **Settings → Providers → Bash commands** to count that spend (it is off by default). Each bash command then gets a local endpoint and a workspace key, and Xum adds the usage to that workspace's Costs tab and to Analytics (source `headless:bash_proxy`).

| Variable | Value |
| - | - |
| `ANTHROPIC_BASE_URL` | `http://127.0.0.1:<port>/anthropic` |
| `OPENAI_BASE_URL` | `http://127.0.0.1:<port>/openai/v1` |
| `ANTHROPIC_API_KEY`, `ANTHROPIC_AUTH_TOKEN`, `OPENAI_API_KEY` | `xum-proxy-…` (one key per workspace) |

* Xum sends the calls with the provider keys from this page. The command never sees the real key.
* Local and Worktree commands reach Xum directly. SSH and Coder commands reach it through a reverse SSH forward (`ssh -R`) that Xum opens to each host, on the host's `127.0.0.1`. The first turn on a host can wait up to 10 seconds for it. The forward uses its own SSH connection without the forwards in your `~/.ssh/config`. A host that refuses forwarding (`AllowTcpForwarding no`) gets no variables, and Xum tries it again after 5 minutes. After a connection failure Xum tries again after 30 seconds. A host that asks for a password or a key passphrase gets no variables: the forward does not prompt, so load the key into `ssh-agent`.
* With `XUM_ALLOW_MULTIPLE_INSTANCES=1`, or while a `xum server` of another process uses the same Xum home, SSH and Coder commands get no variables: a host keeps one forwarded port per Xum home, and two backends would each pick their own. Xum does not detect a desktop app started with `XUM_NO_API_SERVER=1` next to `xum server`.
* Docker and devcontainer commands keep their environment: they have no route to Xum.
* Bash commands in `xum run` and `xum workflow` sessions do not get the variables.
* Commands in untrusted projects, and every command while `XUM_DISABLE_PROJECT_AUTOMATION=1` is set, get no variables: Xum blanks provider keys there, so repo code cannot spend through the proxy. Revoking a project's trust also refuses calls from its commands that are still running.
* A provider without a Xum API key (for example OpenAI with Codex OAuth only) gets no variables. A provider that your [project secrets](/config/project-secrets) configure keeps the secret values. This includes `OPENAI_ORG_ID` and `OPENAI_PROJECT_ID`.
* The proxy allows Messages, Responses, Chat Completions, token counting and model listing. Other paths get 404.
* The port, the keys and the forwarded ports stay the same when Xum restarts, so background commands keep working. Xum keeps them in `~/.xum/bash-ai-proxy.json`. Deleting that file changes all keys. At startup Xum opens the forwards to SSH hosts again, but not to Coder workspaces, because connecting can start a stopped workspace. Their forward comes back with the next agent turn there.
* If the saved port is busy when Xum starts, Xum uses another port for that session and keeps the saved one. Commands started in that session lose access after the next restart.
* A key stops working when its workspace is removed.
* A script that exports its own `ANTHROPIC_API_KEY` but keeps the proxy URL gets 401. Set both the key and the base URL, or neither.
* Agent CLIs that read these variables also go through the proxy. `claude -p` in a bash command then bills the Xum Anthropic key instead of a Claude subscription login.
* CAUTION: on SSH hosts, Xum passes the variables in the remote command line, and any user on a shared host can reach its `127.0.0.1`. Another user who reads the key with `ps` can spend through your provider keys. Turn this off for shared SSH hosts.

## Advanced: Manual Configuration

For advanced options not exposed in the UI, edit `~/.xum/providers.jsonc` directly:

```jsonc theme={null}
{
  "anthropic": {
    "apiKey": "sk-ant-...",
    "baseUrl": "https://api.anthropic.com", // Optional custom endpoint
  },
  "openrouter": {
    "apiKey": "sk-or-v1-...",
    // Provider routing preferences
    "order": ["Cerebras", "Fireworks"],
    "allow_fallbacks": true,
  },
  "xai": {
    "apiKey": "sk-xai-...",
    // Search orchestration settings
    "searchParameters": { "mode": "auto" },
  },
  "bedrock": {
    "region": "us-east-1",
    // Uses AWS credential chain if no explicit credentials
  },
  "ollama": {
    "baseUrl": "http://your-server:11434/api", // Custom Ollama server
  },
}
```

### OpenAI wire format

The built-in `openai` provider supports two wire formats:

| Value | Description |
| - | - |
| `responses` | Uses the OpenAI Responses API. This is the default and supports current OpenAI and Codex features. |
| `chatCompletions` | Uses `/v1/chat/completions`. Use this for OpenAI-API-shaped gateways that do not expose `/v1/responses`, such as Azure Government endpoints. |

OpenAI documents that GPT-6 Astra only supports tool calling through the Responses API, so keep
`responses` when using that model.

Set the format in **Settings → Providers → OpenAI → Wire format**, or in
`~/.xum/providers.jsonc`:

```jsonc theme={null}
{
  "openai": {
    "baseUrl": "https://your-openai-gateway.example/v1",
    "wireFormat": "chatCompletions",
  },
}
```

Like custom providers, a built-in `openai` `baseUrl` without a path gains `/v1` automatically;
add a trailing slash to keep requests at the origin root.

For llama.cpp, vLLM, LM Studio, and multiple local endpoints, prefer named custom
`providerType: "openai-compatible"` providers instead of changing the built-in OpenAI provider.

### Cyber mode (OpenAI Daybreak)

OpenAI's [Daybreak access programs](https://developers.openai.com/api/docs/guides/daybreak)
provide approved access for cybersecurity work; a Responses API request selects one with
`access_programs.cyber`. The request value selects behavior within your approved access and does
not grant access. Turn on
**Settings → Providers → OpenAI → Enable cyber model** (off by default) to add a **Cyber** option
to the reasoning selector next to the model picker, or run "Toggle Cyber Mode" from the Command
Palette. Cyber and Pro are mutually exclusive.

Cyber appears only for models with a documented Daybreak value. GPT-6.1 Sol and GPT-6 Astra send
`daybreak_blue`. OpenAI requires Daybreak Red approval for reduced refusals on these models, but
they still take `daybreak_blue`. Cyber works only with the built-in OpenAI provider's API key on
the `responses` wire format. A custom base URL on the built-in provider also receives the field.
Gateways, custom providers, Codex OAuth, and Chat Completions use standard safeguards.

If OpenAI rejects the program, Xum shows OpenAI's message and does not retry:

| Error code | Meaning |
| - | - |
| `invalid_access_program` | The model needs a different program value. |
| `unsupported_access_program` | OpenAI does not offer Daybreak on this model to you. |
| `access_program_not_enabled` | The API key's project or organization lacks the program. |

### Custom providers

Custom providers let you keep the built-in `openai` and `anthropic` providers for their official
APIs while adding named local endpoints, remote gateways, or proxies. Each custom provider selects
the API format that its endpoint accepts.

Add each endpoint as a top-level provider in `~/.xum/providers.jsonc`:

```jsonc theme={null}
{
  "local-vllm": {
    "providerType": "openai-compatible",
    "displayName": "Local vLLM",
    "baseUrl": "http://localhost:8000/v1",
    "models": ["qwen3-coder"],
  },
}
```

The **API format** selector in Settings defaults to `openai-compatible`, so existing custom
provider setups continue to use Chat Completions.

| `providerType` | SDK adapter | Request endpoint |
| - | - | - |
| `openai-compatible` | `@ai-sdk/openai-compatible` | `/v1/chat/completions` |
| `openai-responses` | `@ai-sdk/openai` Responses API | `/v1/responses` |
| `anthropic-messages` | `@ai-sdk/anthropic` | `/v1/messages` |

| Option | Required | Description |
| - | - | - |
| `providerType` | Yes | API format for this custom provider. Use one of the three values above. |
| `baseUrl` | Yes | API base URL. Xum normalizes it according to `providerType` before the SDK adds the endpoint path. |
| `apiKey` | No | Optional API key. Keyless local servers do not need a placeholder key. |
| `apiKeyFile` | No | Path to a file containing the API key. Supports `~` for home directory. |
| `displayName` | No | Optional UI label shown instead of the provider ID. |
| `models` | No | Discovery list shown in the model picker. You can still type any model ID at runtime. |

Base URL normalization depends on the selected format:

* `openai-compatible` and `openai-responses`: an origin-only URL such as
  `http://localhost:8080` gains `/v1`. Explicit paths remain unchanged. A trailing slash, such as
  `http://localhost:8080/`, keeps requests at the origin root.
* `anthropic-messages`: trailing slashes are removed and `/v1` is appended unless the URL already
  ends in `/v1`. For example, `https://gateway.example/anthropic` becomes
  `https://gateway.example/anthropic/v1`.

Base URLs must not contain a query string or fragment: the SDK appends endpoint paths directly to
the base URL string, so a query would swallow the endpoint path. Pass authentication through
`apiKey` or `apiKeyFile` instead of query parameters.

To configure a single OpenAI-shaped gateway through the built-in `openai` provider instead, see
[OpenAI wire format](#openai-wire-format).

Provider IDs must use lowercase letters, digits, `_`, and `-`, and must start with a letter or digit. They must not collide with
built-in provider names, and cannot contain `.`, `:`, `/`, or whitespace. Good examples are
`local-vllm`, `llama-cpp`, and `lm-studio`.

#### llama.cpp

```jsonc theme={null}
{
  "llama-cpp": {
    "providerType": "openai-compatible",
    "displayName": "llama.cpp",
    "baseUrl": "http://localhost:8080/v1",
    "models": ["qwen3-coder"],
  },
}
```

#### vLLM

```jsonc theme={null}
{
  "local-vllm": {
    "providerType": "openai-compatible",
    "displayName": "Local vLLM",
    "baseUrl": "http://localhost:8000/v1",
    "models": ["qwen3-coder"],
  },
}
```

#### LM Studio

```jsonc theme={null}
{
  "lm-studio": {
    "providerType": "openai-compatible",
    "displayName": "LM Studio",
    "baseUrl": "http://localhost:1234/v1",
    "models": ["local-model"],
  },
}
```

Operational notes:

* `baseUrl` is resolved from the Xum backend process. In desktop mode, this is your local
  machine. In server mode, the endpoint must be reachable from the server.
* Most compatible servers require the `/v1` suffix in the URL. Use the normalization rules
  above when a gateway mounts its API under another path.
* Keep the official providers for provider-specific account features such as Codex OAuth and
  OpenAI service tiers. Custom providers select the request format but do not inherit those
  features.
* Custom providers are direct-only and do not participate in gateway routing.
* You no longer need to set a fake `apiKey` to point Xum at a keyless local server.

### Coder (Login with Coder)

The Coder provider routes requests through a [Coder](https://coder.com) deployment's
AI Gateway (`/api/v2/aibridge`), so usage is authenticated, governed, and audited by the
deployment instead of a personal API key.

Deployment prerequisites (admin-side):

* The OAuth2 provider experiment: start `coderd` with `--experiments=oauth2` (or `CODER_EXPERIMENTS=oauth2`)
* AI Gateway entitlement and `--aibridge-enabled`, with at least one provider configured

To connect:

1. Open **Settings → Providers → Coder**
2. Set the **Deployment URL** (e.g. `https://coder.example.com`)
3. Click **Login with Coder** and approve the request in your browser

Xum registers itself as an OAuth2 client on the deployment (RFC 7591 dynamic client
registration), completes an authorization-code + PKCE flow, and stores the resulting tokens
in `~/.xum/providers.jsonc`. Tokens refresh automatically; **Disconnect** revokes and clears
them.

After login, Xum discovers the deployment's configured AI Gateway providers (instances such as
`anthropic` or `claude-aws-us-east-2`) but does not load their model catalogs, which can hold
thousands of IDs. Each provider's type (anthropic, openai, google, bedrock, openai-compat, ...)
decides the wire protocol Xum speaks to its gateway route. Use **Refresh providers** under
**Model routing** (or run the "Settings: Refresh Coder providers" command) to re-discover
providers without a re-login. A loaded catalog then drops the models of providers that were
removed or changed type.

Press **Load model catalog** (or run the "Settings: Load Coder model catalog" command) to fetch
every provider's catalog. Catalog models are identified as `coder:<provider>/<model>` (e.g.
`coder:anthropic/<model>`, `coder:my-openai/<model>`), and the button then reads **Refresh model
catalog (N)** with the number of loaded IDs. If any provider's catalog request fails (a provider
that serves no catalog is skipped), the first load does not complete and the catalog stays
unloaded until a retry succeeds; once loaded, a failing provider keeps its previous entries.
Catalog loading never adds models to your configured list. Under **Settings →
Models**, select the **Coder** provider, then use the **Model ID** field to search catalog
suggestions or type a custom model ID; typing works even when the catalog is not loaded.
Picking a suggestion adds it immediately; press **Add** or **Enter** to add a typed ID. Nothing
appears in the model picker until you add it. Added models survive catalog refreshes, re-logins
and **Disconnect**. Removing a model only removes it from your list; use the **Route** column to
steer a model away from Coder. Once a catalog is loaded, built-in models route through Coder only
when the catalog lists them; `/model coder:<provider>/<model>` still selects a catalog model for
the current workspace without adding it.

#### Model routing

**Model routing** maps each native provider (Anthropic, OpenAI, Google) to the deployment's AI
Gateway provider that should serve its models. The model picker keeps showing native models such
as `anthropic:<model>`; when route priority picks Coder for one, Xum sends it to the mapped
provider instead (`coder:claude-aws-us-east-2/<model>` on the wire).

* Only providers of the same type are offered: an Anthropic mapping must point at an
  anthropic-type provider, and so on.
* **Default** keeps the built-in behavior: Anthropic and OpenAI models use the gateway provider
  named `anthropic` or `openai`; Google models route through Coder only when mapped.
* A mapping to a provider that is no longer known (or has a different type) shows as "(not
  found)", and that native provider does not route through Coder until you pick another one.
* Mapping does not change whether Coder is used: route priority and per-model **Route**
  overrides still decide that. Explicit `coder:<provider>/<model>` selections always address the
  named provider literally.
* The Coder card's **Routes to** line lists the native providers that currently route through
  Coder, including a mapped Google and excluding a "(not found)" mapping.
* The model picker does not show the route: **Settings → Models** shows it per model, and replies
  from native models sent through Coder are marked "via Coder" in the transcript.
* The deployment's AI Gateway logs attribute Gemini traffic to `openai`, because Xum speaks the
  OpenAI-compatible protocol to google-type providers.
* Mappings are stored as `canonicalRoutes` in the `coder` section of `~/.xum/providers.jsonc`.
  Xum also keeps its own bookkeeping keys there (`discoveredProviders`, `discoveredModels`,
  `staleDiscoveredModels`, `discoveredModelsUnlisted`, `coderCatalogGeneration`); do not edit
  them by hand.

Listing the deployment's providers requires AI Gateway provider read access, which member
accounts may lack. When it is unavailable, loading the catalog probes the default provider
names instead. If your deployment uses custom-named provider instances that Xum cannot
discover (for example to map them under **Model routing**), declare them by hand in
`~/.xum/providers.jsonc`:

```jsonc theme={null}
{
  "coder": {
    "additionalProviders": [{ "name": "llm-proxy", "type": "openai-compat" }],
    "models": ["llm-proxy/llama-3.3-70b"],
  },
}
```

Login works from the desktop app and from a browser connected to a Xum server, including a
remote one: in the desktop app the OAuth callback lands on a loopback listener on your
machine, while in the browser Xum registers a callback URL on the server's own origin
(`/auth/coder/callback`) so the redirect reaches the server directly. A remote server must
be reachable over HTTPS for Coder to accept that callback URL.

### Bedrock Authentication

Bedrock supports multiple authentication methods (tried in order):

1. **Bearer Token** - Single API key via `bearerToken` config or `AWS_BEARER_TOKEN_BEDROCK` env var
2. **Explicit Credentials** - `accessKeyId` + `secretAccessKey` in config
3. **AWS Credential Chain** - Automatic resolution from environment, `~/.aws/credentials`, SSO, EC2/ECS roles

If you're already authenticated with AWS CLI (`aws sso login`), Xum uses those credentials automatically.

### Anthropic Fast mode

Claude Opus 5.5, Opus 5, and Opus 4.8 support Anthropic's [Fast mode](https://platform.claude.com/docs/en/build-with-claude/fast-mode) research preview: up to 2.5× faster output at 2× token pricing. Toggle it from the thinking selector's **Fast mode** row, the command palette (**Toggle Fast Mode**), or its keyboard shortcut. Xum stores the preference as `"speed": "fast"` under `anthropic` in `providers.jsonc`.

Fast mode is only sent on the direct Anthropic API with an account that has Fast mode access. It is hidden for gateway routes (Xum Gateway, OpenRouter, Bedrock, Coder), custom Anthropic-compatible providers, a non-`api.anthropic.com` base URL, and when Anthropic beta features are disabled for ZDR. Costs are priced from the speed Anthropic reports for each response.

### OpenRouter Provider Routing

Control which infrastructure providers handle your requests:

* `order`: Priority list of providers (e.g., `["Cerebras", "Fireworks"]`)
* `allow_fallbacks`: Whether to try other providers if preferred ones are unavailable
* `only` / `ignore`: Restrict or exclude specific providers
* `data_collection`: `"allow"` or `"deny"` for training data policies

See [OpenRouter Provider Routing docs](https://openrouter.ai/docs/features/provider-routing) for details.

### xAI Search Orchestration

Grok models support live web search. Xum enables this by default with `mode: "auto"`. Customize via [`searchParameters`](https://docs.x.ai/docs/resources/search) for regional focus, time filters, or to disable search.

### Model Parameter Overrides

Set per-model defaults for parameters like temperature, token limits, and sampling by adding a
`modelParameters` section under any provider:

```jsonc theme={null}
{
  "anthropic": {
    "apiKey": "sk-ant-...",
    "modelParameters": {
      // Override for a specific model
      "claude-sonnet-4-5": {
        "temperature": 0.7,
        "max_output_tokens": 16384,
      },
      // Wildcard default for all Anthropic models
      "*": {
        "max_output_tokens": 8192,
      },
    },
  },
}
```

#### Supported parameters

| Parameter | Range | Description |
| - | - | - |
| `temperature` | 0–2 | Randomness of responses |
| `top_p` | 0–1 | Nucleus sampling threshold |
| `top_k` | positive integer | Top-K sampling |
| `max_output_tokens` | positive integer | Maximum response length |
| `seed` | integer | Deterministic generation seed |
| `frequency_penalty` | number | Penalize repeated tokens |
| `presence_penalty` | number | Penalize tokens already present |

Any unrecognized key is passed through as a provider-specific option (for example, OpenRouter routing hints).

#### Resolution order

When multiple entries could match, the **first match wins** (no merging across tiers):

1. **Effective model ID** - a dated snapshot like `claude-sonnet-4-5-20250929`
2. **Canonical model ID** - the model you selected, e.g. `claude-sonnet-4-5`
3. **Wildcard `"*"`** - catch-all for that provider

For example, if you configure both `"claude-sonnet-4-5"` and `"*"` with different temperatures,
requesting `claude-sonnet-4-5` uses the specific entry - the wildcard is not merged in.

#### Priority with other settings

For `max_output_tokens` specifically, the priority chain is:

1. Explicit per-message override (from thinking level or UI)
2. `modelParameters` config value
3. Model's built-in default

This means model parameters act as a **default** - they never override explicit per-message choices.

<Warning>
  Some providers require specific parameter values when extended thinking is enabled. For example,
  Anthropic requires `temperature: 1` with thinking. Setting a different temperature in
  `modelParameters` may cause API errors when thinking is active.
</Warning>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.