Skip to content

Get started

Integrity

Define good behavior, connect your agent, follow its trajectory, and improve from evidence.

Integrity connects your agent's actions to the intent behind its task. Define good behavior, connect your existing model client, and follow the task as a continuous trajectory. Use the resulting evidence to improve the instructions, rubrics, and judgment that guide future work.

The Integrity workflow#

Your application keeps its model, tools, and execution loop. Integrity reviews supported proposals before release and records the evidence behind its decisions. The workspace brings that evidence into Runs and Data.

The Integrity workflow · choose a stepFrom intent to evidence

Define what must stay true.

A constitution describes the task’s obligations and limits. Keep authorization tests, respect permission boundaries, and preserve useful work.

Who handles it
Your workspace
What carries forward
Versioned constitution

One task, from intent to outcome#

Throughout these guides, an agent is implementing an authorization change. Its constitution asks it to complete the implementation while preserving authorization tests. After several plausible steps, the agent proposes deleting a failing test. A useful intervention identifies that conflict, preserves the work already completed, and asks for a corrected implementation that keeps the test.

This is a representative example, not a guarantee of any particular judgment. Read the proposal, its supporting evidence, and the corrected continuation together in Runs. Your harness decides when tools execute and when the task is complete.

Connect your application#

Open the existing deployment in your workspace. Use its base URL and platform key alongside your own provider credentials. Persist one session ID for the task and one operation ID before each distinct request. Connect your application shows the Triage SDK, base URL swap and custom endpoint paths. For agent tools you do not build yourself, for example Codex or Claude Code, see Connect your agent fleet.

Turn experience into better judgment#

Review meaningful successes as well as divergences. A custom rubric can reward preserving useful work, solving the task, or explaining uncertainty. Grading makes those expectations testable. Synthetic generation can fill specific gaps in sparse data, and held-out evaluation measures whether a proposed change helps.

Versioned adaptation builds on that evidence where the deployment supports it. The adaptation guide explains Base, customer versions, explicit activation, rollback, and current availability. The security guide explains how credentials, encrypted evidence, and access controls fit together.