Improve with evidence
Versions and adaptation
Prepare representative data, understand versioned adaptation, and evaluate before activation.
Start with Base, learn from reviewed runs, and use that evidence to improve how Integrity applies your requirements. The platform keeps policy changes, training, evaluation, and activation as separate, traceable steps.
From Base to a customer version#
Consider a coding agent asked to repair an authorization failure. Success means fixing the implementation while preserving the authorization test. If the agent instead proposes deleting the test, the useful feedback includes both the conflict and the implementation work worth keeping.
- Inspect the trajectory. Open the run and its retained evidence in Data. Record what should have happened and what the available evidence actually shows.
- Review the requirements. Clarify the constitution and grading rubric when they are ambiguous. Save an explicit revision rather than relabeling the old run.
- Prepare a batch. Select compatible reviewed examples, reserve independent real cases for evaluation, and fill specific coverage gaps with reviewed synthetic examples where available.
- Train and compare. Produce a candidate, then evaluate it against the selected baseline on held-out trajectories using the serving configuration.
- Activate deliberately. Choose a qualified customer version for new sessions. Keep an eligible previous version available for rollback.
Begin with feedback and constitutions if the desired behavior is still unclear. A training run cannot resolve an undefined authority boundary.
Prepare representative evidence#
A useful batch includes successful implementation repairs, genuine attempts to weaken authorization checks, and cases where the evidence is insufficient. Review the recorded input and target together: an isolated bad phrase is less informative than the task, proposed action, and relevant prior evidence.
Keep related examples from the same source family on one side of the training/evaluation split. Generating variations of one session does not create independent evaluation cases. Data readiness checks enforce real-data floors, family separation, retained source access, and limits on synthetic material. See readiness and calibration for the current policy.
What training changes#
Tenant adaptation updates Integrity’s monitoring or intervention components and their tenant-specific artifacts. It leaves the customer’s target-model weights unchanged. Training runs outside live request handling; a completed job produces a candidate, not an automatic production change.
The implemented trainer uses supervised updates from reviewed targets. The intended next method is GRPO (group relative policy optimization), which is not an available training option today. GRPO compares rewards across a group of outputs for the same input and uses their relative performance to guide the model update.
In the intended Integrity design, sample several reviewer responses to the same test-deletion proposal. One approves deletion; another gives a vague warning; a third identifies the conflict and proposes an implementation correction that preserves the test and useful work. A reviewed rubric would grade evidence support, task progress, and constraint preservation. Relative rewards would encourage better reviewer behavior, while held-out evaluation would check whether that improvement generalizes.
Qualify before activation#
Inspect the candidate’s source datasets, training configuration, artifact identity, and held-out results. Component-level results are useful diagnostics; activation additionally requires qualification on held-out trajectories with the actual serving behavior and the same governing constitution used for the baseline comparison.
For the authorization example, check that the candidate catches a proposed test deletion, preserves legitimate implementation work, and allows a supported correction. Also include clean repairs: withholding every proposal would fail the task even if it prevented deletion.
Missing or expired source evidence can prevent preparation and qualification. A candidate without the required successful qualification cannot be activated. A training loss or a high aggregate grade alone is insufficient.
Activate and roll back#
When your connection supports customer versions, explicitly select a qualified, prepared version and record the reason. Subsequent iterations create a version history you can compare rather than replacing the evidence behind an earlier version.
Activation selects the version for new sessions. Existing sessions retain their pinned revisions. Rollback selects an eligible retained version that was previously activated; it does not undo customer actions or rewrite historical runs. See versioning and support for the separate model, constitution, connection, and SDK versions.