Improve with evidence
Data readiness and calibration
Check representative coverage, training readiness, and held-out behavior before a version change.
Calibrate Integrity against reviewed task outcomes: useful work should proceed, material divergence should be caught, and missing evidence should remain visible.
Define the behavior you want#
Start with an observable requirement. For an authorization repair, “retain the protected test and correct the implementation” describes both successful progress and the boundary the agent must preserve. “Be safer” does not tell a reviewer which proposal should pass.
Use grading-rubric weights to express relative consequence. Weights range from 1 to 5 and default to 3. They affect the rubric’s score; they do not grant permission to violate a constitutional constraint. A score is also distinct from the runtime’s decision to release or withhold an action. See grading for how uncertain and inapplicable clauses affect scoring.
Measure coverage before volume#
Review coverage across task types, source families, action boundaries, and outcomes. Include clean successes, observed divergence, plausible near misses, and incomplete evidence. Repeated versions of one authorization-test failure can add volume while leaving other workflows unrepresented.
Generation does not repair missing coverage.
Counts are trajectories in this example. Held-out task families stay separate from training and synthetic generation inputs.
The real examples cover successful requests, but lack permission-denied paths and multi-turn drift.
Collect and review the missing real situations. Synthetic variations can then explore known gaps.
Use synthetic generation to target an identified gap. Keep an independent real evaluation set and continue collecting real examples of unfamiliar tasks. Generated variation cannot establish how the system behaves on a population you have not observed.
Understand batch readiness#
The current platform-managed training admission policy uses these operational floors. They prevent obviously sparse batches; they are not a universal sample-size recommendation or a quality guarantee.
| Check | Current requirement |
|---|---|
| Real training records | At least 32 |
| Real held-out records | At least 8 |
| Independent real training families | At least 4 |
| Independent real held-out families | At least 2 |
| Synthetic share | At most 50%; the configured cap may be lower |
| Split and review | No family leakage; synthetic examples are reviewed and training-only |
Readiness also checks the requirements of the component being trained, dataset compatibility, and continued access to source evidence. Synthetic records do not count toward the real-data floors. Passing these checks permits preparation of a training job; production activation still requires held-out trajectory qualification.
Evaluate decisions in context#
- Unnecessary holds: did Integrity interrupt a valid implementation repair, and which evidence or requirement caused the disagreement?
- Missed divergence: did earlier steps accumulate into an attempt to remove the protected authorization test?
- Preserved work: did the correction retain useful implementation changes and the original task constraints?
- Recovery: was the corrected proposal evaluated before release, and did it satisfy the task with the evidence available?
Compare revisions on the same held-out cases and inspect disagreements before changing the constitution, rubric, or model. Keep these changes separate so you can explain the result. Do not interpret an unqualified model score as a calibrated probability; the current trainer records that calibration was not fitted.
Use the current connection’s supported controls rather than retired classifier thresholds. Continue with versions and adaptation for the qualification and activation sequence.