Follow a task
Errors and recovery
Interpret errors and recover original operations without automatic replay.
Handle a typed hold separately from a transport failure. Neither permits an automatic replacement request. The only exceptions are the two codes below that carry automatic_retry_authorized: true; everything else is terminal for that request until you act.
Recognize an Integrity hold#
{"error":{"type":"integrity_hold","code":"session_id_required","message":"Required session identity is missing","automatic_retry_authorized":false}}The example illustrates the shape; the actual message and assigned operation ID depend on the failure. Preserve the code and original identifiers. A stream can deliver an error after HTTP 200.
Respond to the cause#
| Status or condition | Next step |
|---|---|
| 401 / 403 | Check the intended platform/provider credentials and workspace or deployment access. |
| 401 composite_project_key_invalid / composite_provider_credential_invalid / composite_credential_ambiguous | The single composite credential tsk_<project key>:<provider credential> was malformed: the tsk_ half does not match the key pattern, the provider half is empty, has whitespace, is not printable ASCII or is itself a tsk_ key, or contains a second ':'. Fix the value; the gateway never echoes or logs either half. Available after the gateway release of 2026-09-21; see configuration. |
| 400 / 422 | Correct invalid IDs, unsupported protocol fields, or missing explicit limits. |
| 409 | Inspect a conflict, closed session, unavailable setup, or governing transition. |
| 404 | Check the request path. Retired check URLs are not exposed on the current model gateway. |
| 410 | Migrate a retired check on the historical check service, or an unbound legacy model connection. |
| 503 container_not_ready | Wait retry_after_seconds and send the same request with the same operation ID. See below. |
| 429 / other 5xx / timeout | Inspect the original operation; a transient-looking error does not authorize replay. |
Examples include session_id_required, idempotency_key_required, explicit_output_limit_required, deployment_migration_required, and setup_unavailable. Handle unknown codes conservatively and preserve their evidence.
Codes that authorize a retry with backoff#
Two responses say, in the body, that retrying is correct. Both carry automatic_retry_authorized: true and a retry_after_seconds hint. The reference client clamps the hint to between 1 and 10 seconds and stops when its deadline or call budget runs out.
| Code | Where it appears | What it means | Correct response |
|---|---|---|---|
| container_not_ready | HTTP 503 on a POST; body also carries request_forwarded: false, container, startup_status | The gateway container that answered is still starting. Nothing was issued or forwarded. | Wait retry_after_seconds, then re-send the same request with the same operation ID. Another container may answer. Do not treat it as an outcome. |
| review_capacity_exhausted | A completed evaluation with decision hold; policy_decision has capacity_hold: true and reasons: ["review_capacity_exhausted"] | The review could not be admitted within its bounded wait. No verdict was formed and the action is held; nothing was allowed. | Keep the held evaluation as evidence. Wait retry_after_seconds, then evaluate the same proposal again under a new operation ID. If your evaluation budget is spent, the hold stands as the decision. |
A 503 without both automatic_retry_authorized: true and request_forwarded: false is not a readiness refusal. Read the original operation instead of sending again.
Client budget exhaustion#
The reference client raises operation_observation_exhausted_no_retry when its own deadline or HTTP call budget ran out while it was still observing an operation it had issued. The outcome is unknown to the client; the operation may well have completed on the server. The error carries operation_id, last_phase (the last phase the server reported, or unknown), observation_attempts, and exhausted_by (deadline or http_call_budget).
The correct response is a plain authenticated read-back of GET /v1/integrity/operations with the same x-triage-session-id and the original operation ID as Idempotency-Key (integrity.recover in the SDKs), under a fresh budget. Never resubmit the POST: the operation ID is already issued. This is distinct from operation_not_completed_no_retry, where the server positively reported a terminal phase other than completed; that operation is over and there is nothing to read back for.
Hold reasons#
A hold names its reasons in policy_decision.reasons. Reasons are terminal for the proposal unless a replacement accompanies them, in which case the corrected proposal must be evaluated afresh. Reasons you may now see:
| Reason | Plain meaning | What to do |
|---|---|---|
| placeholder_argument_value | A proposed correction used a placeholder instead of a value: <non-empty-idempotency-key>, {{key}}, TODO, ..., or an empty string where the schema requires a non-empty one. It was refused before any model review. | Supply the real value in the argument. Nothing is repaired or substituted for you. |
| citation_not_found_in_named_root | A correction cited text that does not occur exactly once, as an exact string, in the named authority root. One further attempt is made, then the hold stands. | Quote the authority text exactly. The quote is never normalised or case-folded. |
| review_capacity_exhausted | Review capacity was exhausted within the bounded wait; no verdict was formed. | Retry with backoff as described above. |
| correction_outside_owner_scope | The corrected action changed leaves outside the owner-declared mutable_paths, or more leaves than paths were declared. | Declare every path a correction may touch; one declared path per independent requirement. |
| tool_identity_not_mutable | A declared mutable path was outside /arguments/. The tool itself can never change through model output. | Declare argument paths only. |
Recover an uncertain operation#
Read GET /v1/integrity/operations with the original platform key, session ID, and idempotency key. A completed result can be recovered; an in-progress or unknown result needs reconciliation. Changing IDs can create another operation, not finish the first.
Disable provider-client automatic retries. Continuous Integrity controls already make a single attempt. See configuration and operation recovery.
Provide useful evidence#
Share the SDK/API version, deployment identifier, session and operation IDs, timestamp, hold code, and redacted error shape through your support channel. Do not include credentials or raw sensitive payloads.