Skip to content

Connect your model

Connect your application

Call a deployment through the Triage SDK, your provider's SDK with a base URL swap, or your own endpoint.

A deployment can front any application that calls a model: a support assistant, a document pipeline, a backend job or an agent. Your application keeps its provider SDK, its provider key and its model; it sends requests through the deployment's gateway base URL with the deployment key. Every example below is generated from the same source as the connection instructions in Deployments and was run against the published SDKs.

Before you start#

  • Create or open the deployment in Deployments. Copy its key when it is shown and its base URL from the connection instructions; the examples use https://integrity.triage-sec.com/v1.
  • Export the key as TRIAGE_API_KEY and your provider credential under its usual name. Keep both on the server.
  • Every model request carries the key in x-triage-api-key, one x-triage-session-id for the whole task or conversation, a new Idempotency-Key per request, and an explicit output token limit.
  • Turn provider retries off. After a lost response, read the operation back with the same identifiers instead of sending it again; see errors and recovery.

Choose a path#

PathUse it whenWhat changes in your code
Triage SDKYou want the session, request identity and task close handled for you.Install triage-integrity-sdk 0.6.1; pass its base URL and headers to your provider SDK.
Base URL swapYou want no new dependency.Point your provider SDK at the gateway and add three headers.
Custom endpointYour model runs behind your own OpenAI-compatible or Anthropic-compatible endpoint.The deployment forwards to your endpoint; your code is the same as a base URL swap.

Triage SDK#

The Integrity session client supplies the gateway base URL and the identity headers for each request, and finish closes the task explicitly. Your provider SDK still makes the model call. Install the published release:

Shell
pip install triage-integrity-sdk==0.6.1 "openai>=3.14.1"
npm install @triage-integrity/integrity-sdk@0.6.1 openai
import os
import uuid

from openai import DefaultHttpx2Client, OpenAI
from triage_sdk import Integrity

integrity = Integrity(
    api_key=os.environ["TRIAGE_API_KEY"],
    base_url="https://integrity.triage-sec.com/v1",
    session_id="order-4821",  # one ID per task or conversation
)
client = OpenAI(
    api_key=os.environ["OPENAI_API_KEY"],
    base_url=integrity.base_url,
    http_client=DefaultHttpx2Client(follow_redirects=False),
    max_retries=0,
)

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Draft a two-sentence reply to a customer asking where their order is."}],
    max_completion_tokens=512,
    extra_headers=integrity.headers(uuid.uuid4().hex),  # a new ID per request
)
print(response.choices[0].message.content)

# When the task is over and its last review has completed:
# integrity.finish(uuid.uuid4().hex, outcome="completed")

Reuse one Integrity instance, and so one session ID, for every request of a task. Call finish when the task is over. In Observe each request is reviewed after its response, usually within a few minutes; a finish sent before the last review completes is accepted (HTTP 202) and the session closes when that review finishes. Sending the same call again is safe. With the Anthropic SDK, pass integrity.base_url without its /v1 suffix, because that SDK adds it itself:

import os
import uuid

from anthropic import Anthropic, DefaultHttpxClient
from triage_sdk import Integrity

integrity = Integrity(
    api_key=os.environ["TRIAGE_API_KEY"],
    base_url="https://integrity.triage-sec.com/v1",
    session_id="order-4821",  # one ID per task or conversation
)
client = Anthropic(
    api_key=os.environ["ANTHROPIC_API_KEY"],
    base_url=integrity.base_url.removesuffix("/v1"),  # the Anthropic SDK adds /v1 itself
    http_client=DefaultHttpxClient(follow_redirects=False),
    max_retries=0,
)

message = client.messages.create(
    model="claude-sonnet-4-5",
    max_tokens=512,
    messages=[{"role": "user", "content": "Draft a two-sentence reply to a customer asking where their order is."}],
    extra_headers=integrity.headers(uuid.uuid4().hex),  # a new ID per request
)
print(message.content[0].text)

# When the task is over and its last review has completed:
# integrity.finish(uuid.uuid4().hex, outcome="completed")

Base URL swap#

Keep your provider SDK and change two things: its base URL, and the headers it sends.

import os
import uuid

from openai import DefaultHttpx2Client, OpenAI

client = OpenAI(
    api_key=os.environ["OPENAI_API_KEY"],
    base_url="https://integrity.triage-sec.com/v1",
    default_headers={"x-triage-api-key": os.environ["TRIAGE_API_KEY"]},
    http_client=DefaultHttpx2Client(follow_redirects=False),
    max_retries=0,
)

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Draft a two-sentence reply to a customer asking where their order is."}],
    max_completion_tokens=512,
    extra_headers={
        "x-triage-session-id": "order-4821",  # one ID per task or conversation
        "Idempotency-Key": uuid.uuid4().hex,  # a new ID per request
    },
)
print(response.choices[0].message.content)
import os
import uuid

from anthropic import Anthropic, DefaultHttpxClient

client = Anthropic(
    api_key=os.environ["ANTHROPIC_API_KEY"],
    base_url="https://integrity.triage-sec.com",
    default_headers={"x-triage-api-key": os.environ["TRIAGE_API_KEY"]},
    http_client=DefaultHttpxClient(follow_redirects=False),
    max_retries=0,
)

message = client.messages.create(
    model="claude-sonnet-4-5",
    max_tokens=512,
    messages=[{"role": "user", "content": "Draft a two-sentence reply to a customer asking where their order is."}],
    extra_headers={
        "x-triage-session-id": "order-4821",  # one ID per task or conversation
        "Idempotency-Key": uuid.uuid4().hex,  # a new ID per request
    },
)
print(message.content[0].text)

Custom endpoint#

A deployment can forward to your own endpoint instead of a public provider, for example a self-hosted model server or an internal model router. Choose Custom endpoint when you create the deployment, pick the OpenAI-compatible or Anthropic-compatible format, and enter the HTTPS base URL your endpoint serves /chat/completions or /messages under. An existing deployment can be switched under Routing, Override.

Once the route is in place, call the gateway with your endpoint's key and a model it serves:

import os
import uuid

from openai import DefaultHttpx2Client, OpenAI

client = OpenAI(
    api_key=os.environ["ENDPOINT_API_KEY"],
    base_url="https://integrity.triage-sec.com/v1",
    default_headers={"x-triage-api-key": os.environ["TRIAGE_API_KEY"]},
    http_client=DefaultHttpx2Client(follow_redirects=False),
    max_retries=0,
)

response = client.chat.completions.create(
    model="your-model",  # a model your endpoint serves
    messages=[{"role": "user", "content": "Draft a two-sentence reply to a customer asking where their order is."}],
    max_tokens=512,
    extra_headers={
        "x-triage-session-id": "order-4821",  # one ID per task or conversation
        "Idempotency-Key": uuid.uuid4().hex,  # a new ID per request
    },
)
print(response.choices[0].message.content)

The gateway forwards your credential as Authorization: Bearer for the OpenAI-compatible format and as x-api-key for the Anthropic-compatible format.

OpenAI, Anthropic and Azure#

ProviderCredential variableModel fieldNotes
OpenAIOPENAI_API_KEYYour OpenAI modelChat Completions and Responses at the base URL.
AnthropicANTHROPIC_API_KEYYour Anthropic modelMessages; the Anthropic SDK takes the base URL without /v1.
AzureAZURE_API_KEYYour Azure deployment nameUse the resource endpoint, https://<resource>.openai.azure.com. The gateway calls its v1 API and sends your key in the api-key header.
Custom endpointENDPOINT_API_KEYA model your endpoint servesRequires host approval; see above.

The Azure example is the OpenAI SDK with your Azure key and deployment name:

import os
import uuid

from openai import DefaultHttpx2Client, OpenAI

client = OpenAI(
    api_key=os.environ["AZURE_API_KEY"],
    base_url="https://integrity.triage-sec.com/v1",
    default_headers={"x-triage-api-key": os.environ["TRIAGE_API_KEY"]},
    http_client=DefaultHttpx2Client(follow_redirects=False),
    max_retries=0,
)

response = client.chat.completions.create(
    model="your-deployment-name",  # your Azure deployment name
    messages=[{"role": "user", "content": "Draft a two-sentence reply to a customer asking where their order is."}],
    max_completion_tokens=512,
    extra_headers={
        "x-triage-session-id": "order-4821",  # one ID per task or conversation
        "Idempotency-Key": uuid.uuid4().hex,  # a new ID per request
    },
)
print(response.choices[0].message.content)

A deployment set to Automatic accepts both the OpenAI and the Anthropic APIs and forwards each request to the provider it is written for. A pinned deployment accepts only its provider's API.

Streaming#

For Chat Completions, request usage with the stream so token accounting is complete; an endpoint that omits usage from a stream is held. In Observe, tokens are relayed as they arrive and reviewed afterwards; in Enforce, content is released after review, with heartbeats while it waits.

import os
import uuid

from openai import DefaultHttpx2Client, OpenAI

client = OpenAI(
    api_key=os.environ["OPENAI_API_KEY"],
    base_url="https://integrity.triage-sec.com/v1",
    default_headers={"x-triage-api-key": os.environ["TRIAGE_API_KEY"]},
    http_client=DefaultHttpx2Client(follow_redirects=False),
    max_retries=0,
)

stream = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Draft a two-sentence reply to a customer asking where their order is."}],
    max_completion_tokens=512,
    stream=True,
    stream_options={"include_usage": True},
    extra_headers={
        "x-triage-session-id": "order-4821",  # one ID per task or conversation
        "Idempotency-Key": uuid.uuid4().hex,  # a new ID per request
    },
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
print()

Agent tools and frameworks#

Frameworks that wrap a provider SDK, for example the OpenAI Agents SDK or LangChain, connect the same way: give their client the gateway base URL and the headers above. Agent tools you do not build yourself, for example Codex or Claude Code, have their own settings files; see Connect your agent fleet.