Connect your model
Connect your application
Call a deployment through the Triage SDK, your provider's SDK with a base URL swap, or your own endpoint.
A deployment can front any application that calls a model: a support assistant, a document pipeline, a backend job or an agent. Your application keeps its provider SDK, its provider key and its model; it sends requests through the deployment's gateway base URL with the deployment key. Every example below is generated from the same source as the connection instructions in Deployments and was run against the published SDKs.
Before you start#
- Create or open the deployment in Deployments. Copy its key when it is shown and its base URL from the connection instructions; the examples use
https://integrity.triage-sec.com/v1. - Export the key as
TRIAGE_API_KEYand your provider credential under its usual name. Keep both on the server. - Every model request carries the key in
x-triage-api-key, onex-triage-session-idfor the whole task or conversation, a newIdempotency-Keyper request, and an explicit output token limit. - Turn provider retries off. After a lost response, read the operation back with the same identifiers instead of sending it again; see errors and recovery.
Choose a path#
| Path | Use it when | What changes in your code |
|---|---|---|
| Triage SDK | You want the session, request identity and task close handled for you. | Install triage-integrity-sdk 0.6.1; pass its base URL and headers to your provider SDK. |
| Base URL swap | You want no new dependency. | Point your provider SDK at the gateway and add three headers. |
| Custom endpoint | Your model runs behind your own OpenAI-compatible or Anthropic-compatible endpoint. | The deployment forwards to your endpoint; your code is the same as a base URL swap. |
Triage SDK#
The Integrity session client supplies the gateway base URL and the identity headers for each request, and finish closes the task explicitly. Your provider SDK still makes the model call. Install the published release:
pip install triage-integrity-sdk==0.6.1 "openai>=3.14.1"
npm install @triage-integrity/integrity-sdk@0.6.1 openaiimport os
import uuid
from openai import DefaultHttpx2Client, OpenAI
from triage_sdk import Integrity
integrity = Integrity(
api_key=os.environ["TRIAGE_API_KEY"],
base_url="https://integrity.triage-sec.com/v1",
session_id="order-4821", # one ID per task or conversation
)
client = OpenAI(
api_key=os.environ["OPENAI_API_KEY"],
base_url=integrity.base_url,
http_client=DefaultHttpx2Client(follow_redirects=False),
max_retries=0,
)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Draft a two-sentence reply to a customer asking where their order is."}],
max_completion_tokens=512,
extra_headers=integrity.headers(uuid.uuid4().hex), # a new ID per request
)
print(response.choices[0].message.content)
# When the task is over and its last review has completed:
# integrity.finish(uuid.uuid4().hex, outcome="completed")Reuse one Integrity instance, and so one session ID, for every request of a task. Call finish when the task is over. In Observe each request is reviewed after its response, usually within a few minutes; a finish sent before the last review completes is accepted (HTTP 202) and the session closes when that review finishes. Sending the same call again is safe. With the Anthropic SDK, pass integrity.base_url without its /v1 suffix, because that SDK adds it itself:
import os
import uuid
from anthropic import Anthropic, DefaultHttpxClient
from triage_sdk import Integrity
integrity = Integrity(
api_key=os.environ["TRIAGE_API_KEY"],
base_url="https://integrity.triage-sec.com/v1",
session_id="order-4821", # one ID per task or conversation
)
client = Anthropic(
api_key=os.environ["ANTHROPIC_API_KEY"],
base_url=integrity.base_url.removesuffix("/v1"), # the Anthropic SDK adds /v1 itself
http_client=DefaultHttpxClient(follow_redirects=False),
max_retries=0,
)
message = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=512,
messages=[{"role": "user", "content": "Draft a two-sentence reply to a customer asking where their order is."}],
extra_headers=integrity.headers(uuid.uuid4().hex), # a new ID per request
)
print(message.content[0].text)
# When the task is over and its last review has completed:
# integrity.finish(uuid.uuid4().hex, outcome="completed")Base URL swap#
Keep your provider SDK and change two things: its base URL, and the headers it sends.
import os
import uuid
from openai import DefaultHttpx2Client, OpenAI
client = OpenAI(
api_key=os.environ["OPENAI_API_KEY"],
base_url="https://integrity.triage-sec.com/v1",
default_headers={"x-triage-api-key": os.environ["TRIAGE_API_KEY"]},
http_client=DefaultHttpx2Client(follow_redirects=False),
max_retries=0,
)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Draft a two-sentence reply to a customer asking where their order is."}],
max_completion_tokens=512,
extra_headers={
"x-triage-session-id": "order-4821", # one ID per task or conversation
"Idempotency-Key": uuid.uuid4().hex, # a new ID per request
},
)
print(response.choices[0].message.content)import os
import uuid
from anthropic import Anthropic, DefaultHttpxClient
client = Anthropic(
api_key=os.environ["ANTHROPIC_API_KEY"],
base_url="https://integrity.triage-sec.com",
default_headers={"x-triage-api-key": os.environ["TRIAGE_API_KEY"]},
http_client=DefaultHttpxClient(follow_redirects=False),
max_retries=0,
)
message = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=512,
messages=[{"role": "user", "content": "Draft a two-sentence reply to a customer asking where their order is."}],
extra_headers={
"x-triage-session-id": "order-4821", # one ID per task or conversation
"Idempotency-Key": uuid.uuid4().hex, # a new ID per request
},
)
print(message.content[0].text)Custom endpoint#
A deployment can forward to your own endpoint instead of a public provider, for example a self-hosted model server or an internal model router. Choose Custom endpoint when you create the deployment, pick the OpenAI-compatible or Anthropic-compatible format, and enter the HTTPS base URL your endpoint serves /chat/completions or /messages under. An existing deployment can be switched under Routing, Override.
Once the route is in place, call the gateway with your endpoint's key and a model it serves:
import os
import uuid
from openai import DefaultHttpx2Client, OpenAI
client = OpenAI(
api_key=os.environ["ENDPOINT_API_KEY"],
base_url="https://integrity.triage-sec.com/v1",
default_headers={"x-triage-api-key": os.environ["TRIAGE_API_KEY"]},
http_client=DefaultHttpx2Client(follow_redirects=False),
max_retries=0,
)
response = client.chat.completions.create(
model="your-model", # a model your endpoint serves
messages=[{"role": "user", "content": "Draft a two-sentence reply to a customer asking where their order is."}],
max_tokens=512,
extra_headers={
"x-triage-session-id": "order-4821", # one ID per task or conversation
"Idempotency-Key": uuid.uuid4().hex, # a new ID per request
},
)
print(response.choices[0].message.content)The gateway forwards your credential as Authorization: Bearer for the OpenAI-compatible format and as x-api-key for the Anthropic-compatible format.
OpenAI, Anthropic and Azure#
| Provider | Credential variable | Model field | Notes |
|---|---|---|---|
| OpenAI | OPENAI_API_KEY | Your OpenAI model | Chat Completions and Responses at the base URL. |
| Anthropic | ANTHROPIC_API_KEY | Your Anthropic model | Messages; the Anthropic SDK takes the base URL without /v1. |
| Azure | AZURE_API_KEY | Your Azure deployment name | Use the resource endpoint, https://<resource>.openai.azure.com. The gateway calls its v1 API and sends your key in the api-key header. |
| Custom endpoint | ENDPOINT_API_KEY | A model your endpoint serves | Requires host approval; see above. |
The Azure example is the OpenAI SDK with your Azure key and deployment name:
import os
import uuid
from openai import DefaultHttpx2Client, OpenAI
client = OpenAI(
api_key=os.environ["AZURE_API_KEY"],
base_url="https://integrity.triage-sec.com/v1",
default_headers={"x-triage-api-key": os.environ["TRIAGE_API_KEY"]},
http_client=DefaultHttpx2Client(follow_redirects=False),
max_retries=0,
)
response = client.chat.completions.create(
model="your-deployment-name", # your Azure deployment name
messages=[{"role": "user", "content": "Draft a two-sentence reply to a customer asking where their order is."}],
max_completion_tokens=512,
extra_headers={
"x-triage-session-id": "order-4821", # one ID per task or conversation
"Idempotency-Key": uuid.uuid4().hex, # a new ID per request
},
)
print(response.choices[0].message.content)A deployment set to Automatic accepts both the OpenAI and the Anthropic APIs and forwards each request to the provider it is written for. A pinned deployment accepts only its provider's API.
Streaming#
For Chat Completions, request usage with the stream so token accounting is complete; an endpoint that omits usage from a stream is held. In Observe, tokens are relayed as they arrive and reviewed afterwards; in Enforce, content is released after review, with heartbeats while it waits.
import os
import uuid
from openai import DefaultHttpx2Client, OpenAI
client = OpenAI(
api_key=os.environ["OPENAI_API_KEY"],
base_url="https://integrity.triage-sec.com/v1",
default_headers={"x-triage-api-key": os.environ["TRIAGE_API_KEY"]},
http_client=DefaultHttpx2Client(follow_redirects=False),
max_retries=0,
)
stream = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Draft a two-sentence reply to a customer asking where their order is."}],
max_completion_tokens=512,
stream=True,
stream_options={"include_usage": True},
extra_headers={
"x-triage-session-id": "order-4821", # one ID per task or conversation
"Idempotency-Key": uuid.uuid4().hex, # a new ID per request
},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
print()Agent tools and frameworks#
Frameworks that wrap a provider SDK, for example the OpenAI Agents SDK or LangChain, connect the same way: give their client the gateway base URL and the headers above. Agent tools you do not build yourself, for example Codex or Claude Code, have their own settings files; see Connect your agent fleet.