Skip to content

Operations

Performance and usage

Measure latency, review overhead, holds, and known versus unknown usage.

Measure the whole operation: customer inference, Integrity review, any permitted corrective continuation, and response release. A healthy endpoint or one clean example is not a latency or quality guarantee.

Measure latency at the right boundary#

Record the time from dispatch to the completed released or held result. Separate initial input-contract validation, customer-model inference, action/output review, any corrective continuation, and response release when the run records those stages. A recorded span has its own start, end, offset, and duration; parallel spans can overlap, so summing them is not end-to-end latency. For streaming, distinguish heartbeats from released content and full protocol completion.

Gateway and SDK-check latency distributions are separate populations. Read their sample counts and p50, p95, and p99 together; absent percentiles are unavailable measurements, not zero latency. Use comparable workloads, configured budgets, and model settings when comparing deployments. Report the distribution and held or failed operations alongside successful examples.

Preserve known and unknown usage#

Customer-model call and token limits are enforced through the supported request and response contract. Native review has its own call, evaluation, correction, and elapsed-time limits. Native billed token usage and a reliable price derived from it remain unavailable. Some authenticated retained invocation records expose separately named prepared-input or returned-output token counts; those measurements are not billed tokens.

Do not present missing tokens or cost as zero. A correction can consume additional customer inference and review budget. Use only measurements present in the retained operation record.

Costs reports priced calls separately from all recorded calls. A partial total covers only the priced subset; an unavailable amount is not a zero-dollar estimate. Platform-hosted compute activity is excluded from customer-cost totals, and a Triage-covered amount identifies who paid for an included call rather than an additional charge.

The organization-scoped hosted summary reports compute_usage as recorded compute activity. Current records do not retain enough producer and GPU-count provenance to certify a common duration unit: measurement_basis is unverified_producer, gpu_seconds is null, and measured_calls is zero. calls counts ledger entries; reported_calls counts entries with a reported duration value, without certifying its type or units. Read per-component month-to-date counts through the summary’s as-of snapshot. An empty ledger does not prove that no compute ran. These are not fleet utilization, idle capacity, or billable GPU time.

The authorized project-level native activity panel reads retained evidence separately from that ledger summary. Budget-counter increases count observed debits and returned-original settlements, not successful requests. Recognized paired invocation records may supply enqueue-to-terminal wall intervals, prepared-input tokens, returned-output tokens, or separately labeled predictor-wall and joint-input measurements. These are different producer measurements, can overlap, and are not physical GPU time, utilization, gateway-added latency, or a billing basis. The compiled Input-contract calls are outside the paired-record projection. Missing, expired, unsupported, or unreadable evidence stays unavailable; the retained cohort is not an exhaustive history. Overview refreshes while visible; Costs uses its selected project and pricing snapshot, and its other cost-table filters do not filter this project-level panel. These console reads add no customer SDK method and trigger no retraining.

The demo view displays labelled synthetic examples. Organization-wide outcome and performance summaries exclude demo projects; real accounting excludes specifically identified sample usage and cost records. Creating new samples calls no provider and creates no billed usage. Separately incurred real provider or manual-grading charges remain recorded; a demo label does not erase an actual charge.

Evaluate behavior separately#

A successful fixture test verifies the exercised integration boundary. It does not establish real customer-model quality, recovery quality, or an external tool effect. Assess decisions against representative task evidence and retain their governing revision.

See Reading a run and Versions and adaptation.