SOFTWARE RISK INDEX · FOUR WEEKS
Your quality program has eight dimensions. Your agents need a ninth.
A four-week, TMMi-based assessment of how your company builds and ships software — and whether the AI agents it ships have earned production. Eight dimensions are assessed by people. The ninth is produced by the platform, from evidence, and it keeps being produced after we leave.
CTOs and VPs of Engineering at software companies with 20 to 500 engineers.
An organisation at the top of every maturity scale that ships an ungoverned agent still ships the risk. That is the question your quality program is about to be asked, and it has no standard to answer from.
NINE DIMENSIONS, ONE INSTRUMENT
| 1 | Defect trends | Assessed by people, against TMMi |
| 2 | SDLC and release process | Assessed by people, against TMMi |
| 3 | Test automation maturity | Assessed by people, against TMMi |
| 4 | Performance and scalability readiness | Assessed by people, against TMMi |
| 5 | Security and quality engineering | Assessed by people, against TMMi |
| 6 | Engineering workflow efficiency | Assessed by people, against TMMi |
| 7 | DevOps and CI/CD maturity | Assessed by people, against TMMi |
| 8 | Quality governance | Assessed by people, against TMMi |
| 9 | Agentic release readiness | Produced by the platform, from evidence |
DIMENSION 9 · WHAT A 4 MEANS, CRITERION BY CRITERION
Scored 0 to 4 from the release ledger — absent, ad hoc, defined, evidenced, sealed and enforced. A 4 requires a signed decision whose signature verifies against your published key. A criterion the ledger cannot show is reported as unmeasured, and is not averaged in.
- 9.1 Agent identity
- Models, prompts, tools and permissions pinned to an immutable digest, with canonical model names that carry no provider decoration
- 9.2 Capability profile
- All 13 ARC capabilities profiled by a model-graded assessment, the worst acceptable outcome stated, and the profile sealed
- 9.3 Permission scope
- Each tool scoped to explicit actions and an identity, irreversible actions held by a condition, all inside the signed digest
- 9.4 Evals and adversarial evidence
- Evaluation runs and red-team results pinned to the digest rather than attached to a document
- 9.5 Accountable decision
- A named approver answers the six release questions, with conditions and an expiry, and the decision is sealed on the ledger
- 9.6 Enforcement at deploy
- A fail-closed gate admits this digest and has demonstrably denied a candidate it had no authorization for
- 9.7 Revocation and perishability
- The authorization expires, its status is projected, and the gate that reads that status is installed
- 9.8 Runtime drift
- The deployed agent is re-assessed when its spec, model or tools change, from a live runtime signal
- 9.9 Verifiable evidence
- The decision envelope verifies offline against the published issuer key
- 9.10 Framework mapping
- One sealed run maps to ISO 42001, EU AI Act, NIST AI RMF and AIUC-1 readiness, as evidence and never as certification
FOUR WEEKS
Week 1 · Structured discovery
Interviews with engineering leadership, QA, product and the platform or AI lead — how software and agents actually ship, set against how the process document says they do. An inventory of every agent in or near production: its stack, its owner, and its gate if it has one.
Week 2 · Indexing and scoring
Dimensions 1 to 8 scored against TMMi. Each inventoried agent is brought onto the release ledger and scored on the ten criteria of Dimension 9 from what the ledger can show — and marked unmeasured wherever it cannot.
Week 3 · Pressure testing
Automation coverage and the release pipeline, and the two agent failure modes that do the most damage: a deploy path with no gate, and a permission scope with no bound. Demonstrated rather than asserted — a gate is installed in your own pipeline and made to deny a candidate it has no authorization for.
Week 4 · Delivery
An executive scorecard across all nine dimensions, the top 25 risks ranked by business impact, a 30/60/90-day roadmap, and one sealed Evidence Pack per assessed agent that anyone can verify without trusting us.
WHEN DIMENSION 9 COUNTS AS DELIVERED
- Every production agent has at least one sealed Evidence Pack.
- At least one gate is installed in your pipeline, and it has denied a candidate it had no authorization for.
- Your approver has recorded at least one decision on the ledger.
If these are not true at the end of the four weeks, Dimension 9 is reported as unmeasured — not estimated.
TERMS
If the scorecard, the top 25 risks and the roadmap are not delivered within the four weeks, the work continues at no additional consulting cost until they are.
Fixed fee, quoted per engagement.
AFTER THE FOUR WEEKS
After the engagement, Dimension 9 does not stop. Agents change — a new model, a new tool, an expired authorization — and each change is a new subject on the ledger. That is the subscription: the ninth dimension, kept current.