STEP 4 OF 6 · EVALS — 4B prove

Turn non-determinism into evidence.

Step 4 runs in the opposite order to every vendor, on purpose: prove it before you place it. The assignment arrives from Assign — it is below — so the evidence comes before the runtime choice, not after it. What you earn here sets the ceiling the artifact is allowed to print on 4D place + emit. Two dimensions, never one score: “governed” and “good” are different claims, and a single number would hide whichever is weaker.

4A
draft candidate
4B · YOU ARE HERE
prove
4C · YOU ARE HERE
decide
4D
place + emit

4D may not widen a boundary 4C did not grant. If placing the agent needs a permission the decision withheld, that is a new candidate — back to 4A, and the evidence runs again.

TWO DIMENSIONS · NEVER ONE SCORE
Governance — how it is governedNOT MEASURED

Not run yet — draft the manifest on 4.2, then run and seal it here.

This is the dimension that gates autonomy. Run and seal it on your agent's page — your agents →

Accuracy — whether it does the job wellNOT MEASURED

No accuracy run has reported for this agent. Governance evals below prove how it is governed — they do not measure whether it does the job well.

Not built yet. When it lands it reports per-dimension scores with written criteria, tool-selection accuracy, trajectory findings, and the judge's agreement with human labels (κ). Below κ 0.6 the judge's scores are shown but are not load-bearing — an uncalibrated judge is itself an ungoverned agent.

WHAT A SCORE CAN AND CANNOT BUY

Governance evals gate autonomy. Accuracy evals inform you. Neither raises a ceiling on the danger dimension — money and irreversible actions stay human-owned at any score.

A run is estimated until it is sealed to the ledger, and a sealed run decays after 90 days — evidence has a shelf life, and an expired seal is not a verdict.

Continue → 4.2 place + emitBack to the assignment (step 3)
Evals — run + seal · Scalarion