Turn non-determinism into evidence.
Step 4 runs in the opposite order to every vendor, on purpose: prove it before you place it. The assignment arrives from Assign — it is below — so the evidence comes before the runtime choice, not after it. What you earn here sets the ceiling the artifact is allowed to print on 4.2 place + emit. Two dimensions, never one score: “governed” and “good” are different claims, and a single number would hide whichever is weaker.
Not run yet — draft the manifest on 4.2, then run and seal it here.
This is the dimension that gates autonomy. Run and seal it on your agent's page — your agents →
No accuracy run has reported for this agent. Governance evals below prove how it is governed — they do not measure whether it does the job well.
Not built yet. When it lands it reports per-dimension scores with written criteria, tool-selection accuracy, trajectory findings, and the judge's agreement with human labels (κ). Below κ 0.6 the judge's scores are shown but are not load-bearing — an uncalibrated judge is itself an ungoverned agent.
Governance evals gate autonomy. Accuracy evals inform you. Neither raises a ceiling on the danger dimension — money and irreversible actions stay human-owned at any score.
A run is estimated until it is sealed to the ledger, and a sealed run decays after 90 days — evidence has a shelf life, and an expired seal is not a verdict.