Skip to content

Get started

Feedback and constitutions

Build explicit governing expectations from examples, feedback, and custom rubrics.

Use feedback to make the desired behavior clearer. A reviewed example can improve a grading rubric, motivate a constitution revision, or become evidence for adaptation. Each change has its own review and version history.

Start with a concrete disagreement#

An agent is repairing an authorization failure. After several plausible steps, it proposes deleting the failing test. Open the trajectory in Runs and inspect the linked evidence in Data: the original task, protected constraints, relevant tool results, proposed patch, and Integrity’s review.

Write feedback that another reviewer can apply: “Preserve the authorization test and its assertions. Repair the implementation instead. Keep the useful changes already made.” Include the source records that support this judgment. If the test’s protected status or the proposed change is missing, record the gap.

Improvement workflow · choose a stepA review becomes a deliberate change

Start with a specific decision.

Link feedback to the run and evidence: “The authorization test was correctly preserved, but the proposed repair missed the empty-role case.”

Carries forwardAn evidence-linked review

Separate the observed result from your assessment.

Preparation, activation, and rollback depend on your deployment’s supported capabilities.

Connect feedback, rubrics, and authority#

ArtifactPurpose
FeedbackA reviewed judgment about a specific example, with evidence and an explanation.
Grading rubricObservable criteria and examples for evaluating behavior consistently.
ConstitutionThe deployment’s governing task requirements, constraints, and authority.
Adaptation datasetCompatible reviewed examples selected for a supported model update or evaluation.

A grader helps evaluate an example; its output does not grant permissions. Saving a rubric or reviewing a grade does not rewrite a pinned session’s constitution or change its reviewer. Apply governing changes through the deployment’s supported configuration controls.

Turn examples into a requirement#

Pair a positive example with a nearby failure. “Propose an implementation repair that preserves the protected authorization test” asks for useful capability. “Do not remove, skip, or weaken the test to make the suite pass” identifies the divergence to avoid.

In a custom rubric, give the clause a stable identifier, a clear statement, desirable and undesirable examples, and a weight reflecting its consequence. Link to constitutional clauses only when those clauses actually exist. Keep the criterion observable at the selected boundary: a proposed patch can show an intended change, but cannot by itself prove the tests ran successfully.

The rubric assistant can help draft or revise these clauses. Review its wording and examples before saving. It drafts requirements; it does not execute actions or decide which new authority to grant.

Review, evaluate, and revise#

  1. Review the feedback. Inspect the evidence and distinguish a bad proposal from a mistaken grade or an ambiguous requirement.
  2. Evaluate paired cases. Use a grader with a pinned rubric revision to compare the valid implementation repair, test deletion, and a case with incomplete evidence.
  3. Save a deliberate revision. Update the rubric, constitution, or both when the review supports it. Preserve the earlier revision and its records.
  4. Prepare compatible evidence. Keep reviewed feedback tied to its source. Use supported data preparation and readiness checks before treating examples as a training batch.

Continue with grading and human review, synthetic coverage for sparse data, or versions and adaptation.