ÉVALUATION ARC
De l’instance jouable — l’agent Refund Resolver que vous pouvez exécuter sur /sample. Pas de l’effectif complet que le template livrera.
CE QUE CET AGENT PEUT FAIRE
- C1Heuristic: plans or manages multi-step goals.
- C3Heuristic: has discretion over more than one tool.
- C4Heuristic: an LLM agent that communicates in natural language.
- C6Heuristic: can publish communications representing the org externally.
- C7Heuristic: can execute value-bearing transactions.
- C12Heuristic: reads, writes, or deletes files/data.
GARDE-FOUS L0 REQUIS
base-llm-l0-trusted · L0
Use only LLMs from verified, trusted model developers.
base-llm-l0-no-train · L0
Obtain legally binding no-training and no-logging agreements from LLM API providers.
base-llm-l0-log-io · L0
Log all LLM inputs and outputs for regular review.
base-llm-l0-preeval · L0
Conduct structured pre-deployment evaluation of candidate LLMs for instruction-following, performance, and safety.
base-tool-l0-mcp-auth · L0
Use only MCP servers with robust authentication in production.
base-tool-l0-sandbox · L0
Sandbox-test all untested MCP servers before production deployment.
base-tool-l0-mcp-trusted · L0
Use only MCP servers from verified, trusted developers.
VERDICT DE DÉPLOIEMENT
PAS ENCORE PRÊT
Not ready to ship yet — 30 required L0 safeguard(s) not yet self-attested; 6 relevant risk(s) not yet addressed.
C’est l’état honnête d’une démo publique, pas un défaut du template : personne n’a attesté de garde-fou sur cette instance partagée, donc la porte ne peut pas s’ouvrir. Prenez-en une copie, attestez les garde-fous pour votre déploiement, et le même moteur rejoue la porte sur votre configuration.
6 risques pertinents au seuil par défaut (impact ≥ 3 et vraisemblance ≥ 3). Heuristic assessment (no model configured): scores are conservative estimates from deployment context and capability class. Set ANTHROPIC_API_KEY for a full model-graded risk register. Combinatorial capability pairings, if any, are flagged above.
Statut d’application : self-attested (declared, not independently verified).
Profilé par mots-clés (repli déterministe) ; contrôles et verdict dérivés du code dans les deux cas. This is a technical risk assessment produced with the Agentic Risk & Capability (ARC) Framework to support a deployment decision — audit-ready evidence, not a guarantee of safety, a certification, or a substitute for your organization's own security review.