SOFTWARE RISK INDEX · FOUR WEEKS

Your quality program has eight dimensions. Your agents need a ninth.

A four-week, TMMi-based assessment of how your company builds and ships software — and whether the AI agents it ships have earned production. Eight dimensions are assessed by people. The ninth is produced by the platform, from evidence, and it keeps being produced after we leave.

CTOs and VPs of Engineering at software companies with 20 to 500 engineers.

An organisation at the top of every maturity scale that ships an ungoverned agent still ships the risk. That is the question your quality program is about to be asked, and it has no standard to answer from.

NINE DIMENSIONS, ONE INSTRUMENT

1Defect trendsAssessed by people, against TMMi
2SDLC and release processAssessed by people, against TMMi
3Test automation maturityAssessed by people, against TMMi
4Performance and scalability readinessAssessed by people, against TMMi
5Security and quality engineeringAssessed by people, against TMMi
6Engineering workflow efficiencyAssessed by people, against TMMi
7DevOps and CI/CD maturityAssessed by people, against TMMi
8Quality governanceAssessed by people, against TMMi
9Agentic release readinessProduced by the platform, from evidence

DIMENSION 9 · WHAT A 4 MEANS, CRITERION BY CRITERION

Scored 0 to 4 from the release ledger — absent, ad hoc, defined, evidenced, sealed and enforced. A 4 requires a signed decision whose signature verifies against your published key. A criterion the ledger cannot show is reported as unmeasured, and is not averaged in.

9.1 Agent identity
Models, prompts, tools and permissions pinned to an immutable digest, with canonical model names that carry no provider decoration
9.2 Capability profile
All 13 ARC capabilities profiled by a model-graded assessment, the worst acceptable outcome stated, and the profile sealed
9.3 Permission scope
Each tool scoped to explicit actions and an identity, irreversible actions held by a condition, all inside the signed digest
9.4 Evals and adversarial evidence
Evaluation runs and red-team results pinned to the digest rather than attached to a document
9.5 Accountable decision
A named approver answers the six release questions, with conditions and an expiry, and the decision is sealed on the ledger
9.6 Enforcement at deploy
A fail-closed gate admits this digest and has demonstrably denied a candidate it had no authorization for
9.7 Revocation and perishability
The authorization expires, its status is projected, and the gate that reads that status is installed
9.8 Runtime drift
The deployed agent is re-assessed when its spec, model or tools change, from a live runtime signal
9.9 Verifiable evidence
The decision envelope verifies offline against the published issuer key
9.10 Framework mapping
One sealed run maps to ISO 42001, EU AI Act, NIST AI RMF and AIUC-1 readiness, as evidence and never as certification

FOUR WEEKS

  1. Week 1 · Structured discovery

    Interviews with engineering leadership, QA, product and the platform or AI lead — how software and agents actually ship, set against how the process document says they do. An inventory of every agent in or near production: its stack, its owner, and its gate if it has one.

  2. Week 2 · Indexing and scoring

    Dimensions 1 to 8 scored against TMMi. Each inventoried agent is brought onto the release ledger and scored on the ten criteria of Dimension 9 from what the ledger can show — and marked unmeasured wherever it cannot.

  3. Week 3 · Pressure testing

    Automation coverage and the release pipeline, and the two agent failure modes that do the most damage: a deploy path with no gate, and a permission scope with no bound. Demonstrated rather than asserted — a gate is installed in your own pipeline and made to deny a candidate it has no authorization for.

  4. Week 4 · Delivery

    An executive scorecard across all nine dimensions, the top 25 risks ranked by business impact, a 30/60/90-day roadmap, and one sealed Evidence Pack per assessed agent that anyone can verify without trusting us.

WHEN DIMENSION 9 COUNTS AS DELIVERED

  • Every production agent has at least one sealed Evidence Pack.
  • At least one gate is installed in your pipeline, and it has denied a candidate it had no authorization for.
  • Your approver has recorded at least one decision on the ledger.

If these are not true at the end of the four weeks, Dimension 9 is reported as unmeasured — not estimated.

TERMS

If the scorecard, the top 25 risks and the roadmap are not delivered within the four weeks, the work continues at no additional consulting cost until they are.

Fixed fee, quoted per engagement.

AFTER THE FOUR WEEKS

After the engagement, Dimension 9 does not stop. Agents change — a new model, a new tool, an expired authorization — and each change is a new subject on the ledger. That is the subscription: the ninth dimension, kept current.

Software Risk Index · Scalarion