Sample agent · live demo, no login. Your build stays private.

Resolves refund tickets end-to-end — and won't blow up your Stripe account.

Refund ResolverReads the refund request and the order history, checks it against refund policy, issues the refund, and replies to the customer — escalating anything outside policy for a human to approve.

Watch it run

Simulated run on stubbed tools — no real Stripe/CRM is touched. The governance is real.

Hit play to watch the agent resolve a refund — and catch the one that would have cost you.

Before it acts — we proved how it behaves
EVALS · TURN NON-DETERMINISM INTO EVIDENCE

Deterministic checks that the Policy Enforcement Point holds, denies, or allows each scenario as designed — a starter golden-set from your manifest (governance, not task accuracy). Sealing flips confidence est. → verified.

100%8/8 scenarios passATTESTED — ESTIMATED RUN
dangerous actions heldmust hold / deny at the gatePASS
fail-closed (undeclared)must hold / deny at the gatePASS
reversible writesmust allow the safe pathPASS
safe readsmust allow the safe pathPASS
PER-STEP CONFIDENCE → GATES AUTONOMY (EST.)
dangerous actions held
100% (2/2)
fail-closed (undeclared)
100% (1/1)
reversible writes
100% (1/1)
safe reads
100% (4/4)
SAFE TO ACT / ROUTE TO HUMAN
dangerous actions heldROUTE TO HUMAN
fail-closed (undeclared)ROUTE TO HUMAN
reversible writesROUTE TO HUMAN
safe readsROUTE TO HUMAN

Estimated evals only — build your own agent to seal & verify this run.

Governance evals — deterministic checks that the Policy Enforcement Point holds, denies, or allows each scenario as designed, on a starter golden-set derived from your manifest (synthetic). This proves how the agent is GOVERNED, not task accuracy — connect your historical labelled cases for accuracy evals. Confidence is an estimate until an eval run is sealed to the ledger. A readiness signal, not a guarantee.

Deploy it — it earns autonomy as it proves itself
AUTONOMY LADDER — CEILING IS HARD-GATED BY EVALS
EARNED: SUGGEST

Estimated (unsealed) evals → ceiling capped at SUGGEST. Seal the run to earn Approve.

1ShadowRuns read-only beside your process. Sees everything, touches nothing.EARNED
2SuggestDrafts actions and replies; a human sends every one.CURRENT CEILING
3ApproveActs on routine cases; holds anything below the confidence gates for approval.LOCKED
4AutoFully autonomous within scoped, reversible, non-danger actions only.CLOSED — DANGER DIM
Danger dimension: a money / irreversible action is never fully autonomous, at any eval score. Nothing runs on an estimate.
INTEGRATION MAP — YOUR SYSTEMS, NO MIGRATION
Striperead, spend · irreversiblePAYMENTS
Salesforceread, writeCRM
Gmailread, send · irreversibleCOMMS
Refund PolicyreadPAYMENTS
KPI / ROI monitor: not live until connected. We don’t chart fiction.
AUDIT TRAIL · SEALED TO THE LEDGER

No sealed events yet — run a sandboxed test or a governance eval to start the tamper-evident trail.

Autonomy is GATED by your governance evals — the eval ceiling is the maximum; a dimension can never run above what its evals earned, and a dangerous action is never fully autonomous. Live monitoring, SLAs, and ROI activate only when the agent runs on a connected system (v1) — until then this canvas shows the plan and the real audit trail, never fabricated telemetry.

Measure it — the before, the target, the actuals
OUTCOME · BASELINE → TARGET → ACTUALS
BaselineTargetActual
Volume200/wk260/wknot live↑ 30% (target)
Time per case12 min3 minnot live↓ 75% (target)
Error rate8%2%not live↓ 75% (target)
Time to proven value: not live — activates when the agent runs on a connected system

Baseline and target are declared estimates captured at build — the 'before' and the goal, not measured. Actuals stay blank until a connected system reports them. We don't chart fiction.

Operate it — the runbook, from the manifest
RUNBOOK · FOR THE PERSON COVERING TUESDAYRefund Resolver
What it does
  • Resolve customer refund tickets end-to-end
What it can touch
  • Stripe — read, spend · irreversible
  • Salesforce — read, write
  • Gmail — read, send · irreversible
  • Refund Policy — read
When it holds (the gate)
  • Spend over $500.00 per run — held. An undeclared or invalid amount is also held (fail-closed).
Who approves · autonomy
  • Autonomy ceiling: Suggest — capped here until a sealed eval run earns more.
  • A human approves everything held above; nothing danger-class runs unattended.
How to pause · escalate
  • Pause it: disable the agent in your workspace, or revoke its scoped credentials at the connected system — it stops acting at once.
  • Escalation: anything outside the worst-acceptable-outcome (“issuing a refund over $500 without a human approving it”) routes to a human.
  • Re-seal cadence: a sealed eval run stays valid 90 days — after that the autonomy ceiling decays to the unsealed cap until you re-run & re-seal the evals.
  • The tamper-evident ledger records every held and executed action — read it for what happened and why.

Operational runbook derived from the declared manifest — self-attested (what the agent is configured to do), not a record of live operation. Regenerate it whenever the manifest changes.

What it moves once it's live
Time saved~9 hrs / weekrefund tickets handled without a human touching them
Risk held~3 / monthout-of-policy refunds caught before the money moves
Cost$0.05 / ticketvs. ~$6 to resolve one by hand

Illustrative — your numbers once it runs on your Stripe. Not a customer result.

Build your own — free, no login

Every agent you describe is governed exactly like this — from the first message. Watch yours assemble, eval, and earn its autonomy.

Start from scratch
See a governed agent run — Scalarion