Resolves refund tickets end-to-end — and won't blow up your Stripe account.
Refund Resolver — Reads the refund request and the order history, checks it against refund policy, issues the refund, and replies to the customer — escalating anything outside policy for a human to approve.
Simulated run on stubbed tools — no real Stripe/CRM is touched. The governance is real.
Hit play to watch the agent resolve a refund — and catch the one that would have cost you.
Deterministic checks that the Policy Enforcement Point holds, denies, or allows each scenario as designed — a starter golden-set from your manifest (governance, not task accuracy). Sealing flips confidence est. → verified.
Estimated evals only — build your own agent to seal & verify this run.
Governance evals — deterministic checks that the Policy Enforcement Point holds, denies, or allows each scenario as designed, on a starter golden-set derived from your manifest (synthetic). This proves how the agent is GOVERNED, not task accuracy — connect your historical labelled cases for accuracy evals. Confidence is an estimate until an eval run is sealed to the ledger. A readiness signal, not a guarantee.
Estimated (unsealed) evals → ceiling capped at SUGGEST. Seal the run to earn Approve.
No sealed events yet — run a sandboxed test or a governance eval to start the tamper-evident trail.
Autonomy is GATED by your governance evals — the eval ceiling is the maximum; a dimension can never run above what its evals earned, and a dangerous action is never fully autonomous. Live monitoring, SLAs, and ROI activate only when the agent runs on a connected system (v1) — until then this canvas shows the plan and the real audit trail, never fabricated telemetry.
| Baseline | Target | Actual | ||
|---|---|---|---|---|
| Volume | 200/wk | 260/wk | not live | ↑ 30% (target) |
| Time per case | 12 min | 3 min | not live | ↓ 75% (target) |
| Error rate | 8% | 2% | not live | ↓ 75% (target) |
Baseline and target are declared estimates captured at build — the 'before' and the goal, not measured. Actuals stay blank until a connected system reports them. We don't chart fiction.
- Resolve customer refund tickets end-to-end
- Stripe — read, spend · irreversible
- Salesforce — read, write
- Gmail — read, send · irreversible
- Refund Policy — read
- Spend over $500.00 per run — held. An undeclared or invalid amount is also held (fail-closed).
- Autonomy ceiling: Suggest — capped here until a sealed eval run earns more.
- A human approves everything held above; nothing danger-class runs unattended.
- Pause it: disable the agent in your workspace, or revoke its scoped credentials at the connected system — it stops acting at once.
- Escalation: anything outside the worst-acceptable-outcome (“issuing a refund over $500 without a human approving it”) routes to a human.
- Re-seal cadence: a sealed eval run stays valid 90 days — after that the autonomy ceiling decays to the unsealed cap until you re-run & re-seal the evals.
- The tamper-evident ledger records every held and executed action — read it for what happened and why.
Operational runbook derived from the declared manifest — self-attested (what the agent is configured to do), not a record of live operation. Regenerate it whenever the manifest changes.
Illustrative — your numbers once it runs on your Stripe. Not a customer result.
Build your own — free, no login
Every agent you describe is governed exactly like this — from the first message. Watch yours assemble, eval, and earn its autonomy.