Connect your model
Connect your agent fleet
Point Codex, Claude Code, and compatible clients at the Integrity gateway with your own provider credential.
The Integrity gateway at https://integrity.triage-sec.com sits between your agents and your model provider. Every request carries two credentials: your own provider key, which the gateway forwards to the provider your deployment is configured for and never stores, and an Integrity deployment key, which selects the policy that governs the request. Integrity reviews the proposed actions and final output of each request before releasing the provider's response to the agent, and holds anything it cannot release.
Connecting a tool means changing its model base URL and adding the deployment key. The sections below give the exact settings for OpenAI Codex, Anthropic Claude Code, Cursor, and OpenAI- or Anthropic-compatible clients, then the organization-wide rollout, what a held step looks like to the person using the tool, and the current limits.
What your team will see during the trial#
A new deployment starts in Observe mode, and a trial deployment runs in Observe mode for its first week. In Observe mode Integrity reviews every action the same way it would under enforcement and records the verdict it would have given, but it never blocks or alters the agent's traffic: every request is forwarded to the provider and every response is returned to the tool unchanged.
- On Runs. Each session appears live as its tool works. Every turn shows the verdict Integrity would have given, the stage it was judged at, the reasoning behind it, and its timing. Each verdict is labelled as recorded, not enforced, so a reviewer never mistakes a would-have-held step for one that was held.
- For the engineers using the tools. Nothing changes. Codex and Claude Code behave exactly as they did before the base URL changed: no held steps, no messages from Integrity, no extra approvals.
- Latency. In Observe mode the review runs alongside the request rather than gating it, so the tools are not made to wait on Integrity's decision.
- Switching to Enforce. When you are ready, an Authorizer switches the deployment to Enforce from its panel in the deployment inspector, with a reason and a confirmation that names the deployment. The change takes effect on the next request the gateway admits; nothing has to be pushed to devices.
- What Enforce changes. The same verdicts now act on traffic: a step Integrity cannot release is held, and the hold is delivered to the agent in the tool's own protocol exactly as described in What a held step looks like. Returning to Observe later is a reviewed, expiring exception.
Before you roll out#
A fleet rollout depends on four things being true of the deployment you point the tools at. Confirm them in the deployment inspector and with your Integrity contact before pushing any settings.
| Requirement | Why it matters | If it is missing |
|---|---|---|
| A deployment with a revealed key | The key (tsk_...) identifies the deployment on every request. See Connect a deployment. | HTTP 401 Valid Triage API key required. |
| The provider matches the tool's wire protocol | By default the gateway infers the provider from each request. A deployment with a manual routing override pins one upstream provider instead: Codex and OpenAI-compatible clients then need an OpenAI-wire override (openai, openai-chatgpt, openai-compatible, or azure-openai), and Claude Code and Anthropic-compatible clients an Anthropic-wire one (anthropic or anthropic-compatible). | HTTP 400: the upstream provider does not serve that endpoint. |
| Fleet session inference enabled on the deployment | Codex and Claude Code cannot mint the per-task x-triage-session-id and per-request Idempotency-Key that the session contract requires. Integrity enables fleet session inference on the deployment so the session is named from identifiers the tools already send: the Codex thread ID and the Claude Code session ID. Retransmissions of an identical request replay the recorded decision. | HTTP 422 session_id_required on every request. |
| A default output limit (Codex only) | Codex sends no max_output_tokens. The deployment must declare the output limit that applies when a request carries none; nothing is guessed. | Hold with code explicit_output_limit_required. |
What every request carries#
| Header | Value | Notes |
|---|---|---|
x-triage-api-key | The deployment key, tsk_... | Required. On POST /v1/messages only, a tsk_ value sent in x-api-key is also accepted as the deployment key (the Claude Code dual-header form below). It is never forwarded to the provider. |
| Provider credential | Authorization: Bearer <key> for Chat Completions and Responses; x-api-key or Authorization for Messages | Required. Forwarded to the provider the request is routed to, in the header that provider expects. Sending only api-key returns HTTP 401. |
x-triage-provider | A provider family, for example openai or anthropic | Leave it out. An automatic deployment infers the provider from the request and a pinned deployment uses its own route; a value that differs from the route, or openai-chatgpt on an automatic deployment, is refused with HTTP 400 rather than rerouted. |
x-triage-client | codex, claude_code, openai_sdk, anthropic_sdk, triage_sdk, or custom | Optional. Names the calling tool in Runs. Codex and Claude Code are recognized from their own headers when it is absent. |
x-triage-session-id and Idempotency-Key | Your persisted task and request IDs | Required for application code and SDK clients. Codex and Claude Code omit them and rely on fleet session inference. |
OpenAI-wire clients use https://integrity.triage-sec.com/v1 as their base URL. Anthropic clients, including Claude Code, use https://integrity.triage-sec.com without /v1, because they append /v1/messages themselves. Read Authentication for the direct HTTP form.
Codex CLI, IDE, and app#
Codex reads custom model providers only from the user-level ~/.codex/config.toml or a managed layer; a project-local .codex/config.toml cannot set model_provider or model_providers. Codex speaks the Responses protocol; the gateway infers an OpenAI-wire provider from it, and a deployment with a manual routing override must name an OpenAI-wire provider.
# ~/.codex/config.toml (user level). Keep model_provider above the first [table].
model_provider = "triage_integrity"
[model_providers.triage_integrity]
name = "Triage Integrity gateway"
base_url = "https://integrity.triage-sec.com/v1"
wire_api = "responses"
env_key = "OPENAI_API_KEY"
http_headers = { "x-triage-client" = "codex" }
env_http_headers = { "x-triage-api-key" = "TRIAGE_API_KEY" }export OPENAI_API_KEY="<your OpenAI API key>"
export TRIAGE_API_KEY="tsk_..."
codexenv_keynames the environment variable holding your OpenAI API key; Codex sends it asAuthorization: Bearer. This is the supported way to run Codex through the gateway. No provider hint is needed: an automatic deployment infers the provider from the request, and a pinned deployment uses its own route.env_http_headersreads the deployment key fromTRIAGE_API_KEYat startup, so the key never sits in the TOML file.- Codex calls
GET /v1/modelswhen it starts. The gateway authenticates the deployment key and answers with a well-formed model list; Codex then uses its built-in catalog. Choose models by the exact IDs your provider account allows.
Codex signed in with ChatGPT needs a deployment whose routing is set to ChatGPT Codex (Routing, Override, in the deployment's panel). An automatic deployment sends a ChatGPT sign-in to the OpenAI API, where it fails, and a deployment set to ChatGPT Codex cannot serve Codex with an API key, so use one deployment per sign-in method. Replace env_key with requires_openai_auth = true. The device also needs network access to chatgpt.com, which Codex contacts directly in this mode.
model_provider = "triage_integrity"
[model_providers.triage_integrity]
name = "Triage Integrity gateway"
base_url = "https://integrity.triage-sec.com/v1"
wire_api = "responses"
requires_openai_auth = true
http_headers = { "x-triage-client" = "codex" }
env_http_headers = { "x-triage-api-key" = "TRIAGE_API_KEY" }Claude Code#
Claude Code uses its standard gateway environment. Set the base URL to the gateway origin without /v1, keep your Anthropic API key as the model credential, and carry the deployment key in a custom header. Signing in with a Claude plan through the gateway is not supported; set ANTHROPIC_API_KEY.
export ANTHROPIC_BASE_URL="https://integrity.triage-sec.com"
export ANTHROPIC_API_KEY="<your Anthropic API key>"
export TRIAGE_API_KEY="tsk_..."
export ANTHROPIC_CUSTOM_HEADERS="x-triage-api-key: ${TRIAGE_API_KEY}
x-triage-client: claude_code"
claudeClaude Code sends ANTHROPIC_API_KEY as x-api-key and ANTHROPIC_AUTH_TOKEN as Authorization: Bearer; an apiKeyHelper value is sent in both. The gateway accepts the provider credential in either header on /v1/messages and forwards it in the header the routed provider expects. ANTHROPIC_CUSTOM_HEADERS takes one Name: Value pair per line and is where x-triage-api-key belongs. ANTHROPIC_AUTH_TOKEN takes effect immediately; ANTHROPIC_API_KEY asks each user to approve the key once in an interactive session before Claude Code uses it.
The older dual-header form, with the deployment key in ANTHROPIC_API_KEY and the provider key in ANTHROPIC_AUTH_TOKEN, is still accepted for compatibility. The gateway recognizes the tsk_ prefix in x-api-key, uses it as the deployment key, and forwards only the bearer credential. Prefer the custom-header form so each key travels in a header named for its purpose.
export ANTHROPIC_BASE_URL="https://integrity.triage-sec.com"
export ANTHROPIC_API_KEY="tsk_..."
export ANTHROPIC_AUTH_TOKEN="<your Anthropic API key>"
claudeCursor#
Cursor accepts an OpenAI API key and an override for the OpenAI base URL, and nothing else: it sends the key as the only credential and cannot attach x-triage-api-key. For Cursor and any tool with one API key field and no custom headers, the gateway accepts a composite key that carries both halves in the one field. The gateway splits it, admits the deployment with the tsk_ half, and forwards only the provider half upstream on that provider's credential header. Neither half is stored or logged.
tsk_<deployment key>:<your OpenAI API key>
Example shape (not real keys):
tsk_9f2c1a...e7b4:sk-proj-Ab3d...9xYzIn Cursor, open Cursor Settings (Cmd+Shift+J on macOS, Ctrl+Shift+J elsewhere) and go to Models > API Keys:
- In OpenAI API Key, paste the composite key
tsk_<deployment key>:<your OpenAI API key>. - Turn on Override OpenAI Base URL and set it to
https://integrity.triage-sec.com/v1. - Click Verify. Cursor sends a Chat Completions request through the gateway; a green check means the deployment key and the provider key were both accepted.
- Enable only OpenAI models in the model list. Cursor speaks the OpenAI wire, so the gateway routes these requests to OpenAI; a Claude model on that wire is refused as ambiguous unless the deployment has a manual routing override.
The composite is strict. The tsk_ half must be a deployment key, the provider half must be a single printable token, exactly one : separates them, and the composite must be the only credential on the request. Anything else is refused with HTTP 401 and a code naming the problem (composite_project_key_invalid, composite_provider_credential_invalid, composite_credential_ambiguous); the value itself is never echoed. Because Cursor sends requests from its own servers without session identifiers, the deployment needs fleet session inference enabled, as for Codex and Claude Code. Only models used with your own key go through the gateway: Cursor's plan models and its Tab completions always use Cursor's backend and stay outside it; record them in your inventory as such.
GitHub Copilot can reach the gateway only through its bring-your-own-key custom model endpoints, with the composite key in the key field; Copilot plan models use GitHub's own service and stay outside the gateway. The composite key has not yet been verified with Copilot. Gemini CLI is not supported: the gateway does not speak the Gemini API.
OpenAI- and Anthropic-compatible clients#
Any client that lets you set a base URL, an API key, and custom headers can connect. Application code must also persist and send x-triage-session-id and Idempotency-Key; the Python and TypeScript guides show the full session contract. Disable the client's automatic retries so an uncertain operation is recovered by its original ID rather than resent.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["OPENAI_API_KEY"],
base_url="https://integrity.triage-sec.com/v1",
max_retries=0,
default_headers={
"x-triage-api-key": os.environ["TRIAGE_API_KEY"],
},
)A tool that offers only a base URL and a single API key field, with no way to add a header, connects with the composite key: tsk_<deployment key>:<provider key> in the key field, sent as Authorization: Bearer on the OpenAI wire or x-api-key on the Anthropic wire. Leave x-triage-provider out: an automatic deployment infers the provider and a pinned deployment uses its own route, while a hint that does not match the route is refused with HTTP 400.
Roll out to the organization#
Push the same settings to every device from device management or configuration management instead of asking each engineer to edit files. Keep both keys out of the pushed payload: reference environment variables or a credential helper, and deliver the values through your secret distribution.
Codex managed configuration#
Codex applies managed defaults at startup, overriding the user's config.toml and any --config flags. Place the same provider table in /etc/codex/managed_config.toml on Linux and macOS or ~/.codex/managed_config.toml on Windows, or push it on macOS as the base64-encoded config_toml_base64 preference in the com.openai.codex domain through Jamf, Kandji, Fleet, or a similar MDM. Deliver TRIAGE_API_KEY and OPENAI_API_KEY to the device environment separately.
# /etc/codex/managed_config.toml (Linux, macOS) or ~/.codex/managed_config.toml (Windows).
# On macOS, the same text base64-encoded (no wrapping) can be pushed as
# com.openai.codex : config_toml_base64.
model_provider = "triage_integrity"
[model_providers.triage_integrity]
name = "Triage Integrity gateway"
base_url = "https://integrity.triage-sec.com/v1"
wire_api = "responses"
env_key = "OPENAI_API_KEY"
http_headers = { "x-triage-client" = "codex" }
env_http_headers = { "x-triage-api-key" = "TRIAGE_API_KEY" }Claude Code managed settings#
Claude Code reads managed-settings.json from /Library/Application Support/ClaudeCode/ on macOS, /etc/claude-code/ on Linux and WSL, and C:\Program Files\ClaudeCode\ on Windows. The same keys can be delivered as a macOS configuration profile in the com.anthropic.claudecode domain or as the Settings value under HKLM\SOFTWARE\Policies\ClaudeCode. Nothing a user or project sets overrides a managed value.
{
"env": {
"ANTHROPIC_BASE_URL": "https://integrity.triage-sec.com",
"ANTHROPIC_CUSTOM_HEADERS": "x-triage-api-key: tsk_...\nx-triage-client: claude_code",
"CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1"
},
"apiKeyHelper": "/usr/local/libexec/anthropic-provider-key"
}- The
envblock setsANTHROPIC_BASE_URLandANTHROPIC_CUSTOM_HEADERSfor every session; separate header pairs with\ninside the JSON string. - Supply the provider key with an
apiKeyHelpercommand that prints it from the device's secret store, or deliver it asANTHROPIC_AUTH_TOKENin the device environment. Both reach the gateway in a header it accepts. - Server-managed settings from the claude.ai console are not fetched while
ANTHROPIC_BASE_URLpoints at a non-Anthropic host, so use the file or MDM path for a gateway fleet.forceLoginMethodandforceLoginOrgUUIDcannot coexist with a gateway credential.
Deployments and keys#
- One deployment serves one policy. Use one deployment per team, or one per organization, according to how you want to review and enforce. With the default inferred routing, one deployment can serve Codex and Claude Code together; only a deployment with a manual routing override is pinned to one provider family, so an organization running both tools on overrides needs at least two deployments, one on each wire.
- Every member of a deployment shares its key. Sessions in Runs are named by each tool's own thread or session identifier, so one key does not merge two engineers' work into one task.
- Rotate or revoke from the deployment inspector, then push the new value. A revoked key is refused on the next request from every device at once. See Connect a deployment and Bring your own key.
What a held step looks like#
Streaming tools have already received HTTP 200 by the time Integrity decides, so a hold cannot be an HTTP error. By default the gateway completes the stream with a single assistant message that explains the outcome, in the tool's own protocol, and the tool renders it as the model's reply. No provider call was made and no tool ran.
Integrity held this step: it was not sent to the model provider and no tool ran.
Reason: <plain-language reason> (code <hold_code>).
Operation op-... is recorded in Integrity; open the run there for the full evidence.
What you can do: rephrase or narrow the step, or ask an Authorizer to review the run; retrying the same request replays this decision.- A denied step reads
Integrity denied this step: the proposed action violates the deployment's policy and was not executed.and asks the user to change the approach or ask an Authorizer. - A pending step reads
Integrity has not finished reviewing this step; its outcome is pending and nothing has executed.and asks the user to wait, then send the same request again unchanged. - Non-streaming requests receive the hold as a JSON error with the HTTP status of the hold and an
x-triage-integrity-dispositionheader ofheld,denied, orpending.
Tell your engineers what to do: read the reason, narrow or rephrase the step, and continue. Retrying the identical request replays the recorded decision instead of re-running the review, so a retry loop does not change the outcome. The operation ID in the message is the one shown in Runs, where an Authorizer can review the evidence. See Release and holds for the application-side contract.
Limits#
| Boundary | Current behavior |
|---|---|
| Provider families | OpenAI wire: openai, openai-chatgpt, openai-compatible, azure-openai on /v1/chat/completions and /v1/responses. Anthropic wire: anthropic, anthropic-compatible on /v1/messages. Any other provider value is refused with HTTP 400; the gateway never guesses a route or falls back to a default provider. |
| Provider mismatch | By default the provider is inferred from each request and no hint is needed. On a deployment with a manual routing override, an explicit x-triage-provider that differs from the override, or a request whose wire protocol the override does not serve, is refused with HTTP 400. |
| Auxiliary endpoints | POST /v1/responses/compact and POST /v1/messages/count_tokens are refused with HTTP 422 explicit_history_required on an Integrity deployment. Only the model request paths are reviewed and released. |
| Model discovery | GET /v1/models authenticates the deployment key and returns a well-formed list; it does not proxy the provider's catalog except on the ChatGPT-login route. Use the model IDs your provider account allows. |
| Tools that cannot add a header | Cursor and any tool that sends a single API key with no custom header connect with the composite key tsk_<deployment key>:<provider key>. A composite presented next to any other credential, or with a malformed half, is refused with HTTP 401. |
| Hosted agents | A hosted agent with no configurable model base URL, and any browser, shell, or tool traffic that does not pass through the model request, is outside the gateway. |
| Session inference | Fleet session inference recognizes the identifiers Codex and Claude Code send. Other clients must send x-triage-session-id and Idempotency-Key themselves. |
| Streaming | Responses are buffered until release checks finish; the gateway sends protocol heartbeats meanwhile. See Streaming. |
Protocol compatibility does not mean every provider feature is supported. Read Protocols and requests and Limits and FAQ for the request contract, and Errors and recovery for the status codes above.