ARC ASSESSMENT
Of the playable instance — the Refund Resolver agent you can run at /sample. Not of the wider workforce the template will ship.
WHAT THIS AGENT CAN DO
- C1Heuristic: plans or manages multi-step goals.
- C3Heuristic: has discretion over more than one tool.
- C4Heuristic: an LLM agent that communicates in natural language.
- C6Heuristic: can publish communications representing the org externally.
- C7Heuristic: can execute value-bearing transactions.
- C12Heuristic: reads, writes, or deletes files/data.
REQUIRED L0 SAFEGUARDS
base-llm-l0-trusted · L0
Use only LLMs from verified, trusted model developers.
base-llm-l0-no-train · L0
Obtain legally binding no-training and no-logging agreements from LLM API providers.
base-llm-l0-log-io · L0
Log all LLM inputs and outputs for regular review.
base-llm-l0-preeval · L0
Conduct structured pre-deployment evaluation of candidate LLMs for instruction-following, performance, and safety.
base-tool-l0-mcp-auth · L0
Use only MCP servers with robust authentication in production.
base-tool-l0-sandbox · L0
Sandbox-test all untested MCP servers before production deployment.
base-tool-l0-mcp-trusted · L0
Use only MCP servers from verified, trusted developers.
DEPLOY VERDICT
NOT READY YET
Not ready to ship yet — 30 required L0 safeguard(s) not yet self-attested; 6 relevant risk(s) not yet addressed.
That is the honest state of a public demo, not a defect in the template: no one has attested a safeguard on this shared instance, so the gate cannot clear. Take a copy, attest the safeguards for your own deployment, and the same engine re-runs the gate on your configuration.
6 relevant risks at the default threshold (impact ≥ 3 and likelihood ≥ 3). Heuristic assessment (no model configured): scores are conservative estimates from deployment context and capability class. Set ANTHROPIC_API_KEY for a full model-graded risk register. Combinatorial capability pairings, if any, are flagged above.
Enforcement status: self-attested (declared, not independently verified).
Keyword-profiled (deterministic fallback); controls and verdict are code-derived either way. This is a technical risk assessment produced with the Agentic Risk & Capability (ARC) Framework to support a deployment decision — audit-ready evidence, not a guarantee of safety, a certification, or a substitute for your organization's own security review.