Agent Certified
Constructed demonstration ยท Rubric of Methodology v2.0

One agent, scored in the open.

This page scores one AI agent deployment, dimension by dimension, so a reader can see how evidence becomes a number and why the weighted total is not the tier. The deployment is constructed. It is not a real organisation and it is not a certification. As of 17 August 2026, when this page was first published, no organisation had been assessed under this methodology.

StatusConstructed example Published17 August 2026 Weighted total58 of 100 TierCertified

Seven scores, one uneven shape.

The seal is drawn from this deployment's own sub scores. It reaches out where governance and product work are strong and draws in where access, provenance and integration are thin. That inward pull is what decides the tier.

  1. 01

    Trust and Safety

    6 of 10

    Two external red team exercises and tested refusal behaviour. Both used the same scenarios, and production misuse is only sampled.

  2. 02

    Context Integrity

    5 of 10

    A controlled corpus with a named owner. Freshness is asserted, never measured, and no provenance travels with a retrieved passage.

  3. 03

    Distribution Control

    4 of 10

    Authenticated callers and separated environments. No rate limits or spend caps, and blast radius documented for two of six tools.

  4. 04

    Product Maturity

    7 of 10

    Versioned releases, regression on every change, eighteen months of run history. The held out set overlaps the development examples.

  5. 05

    Governance

    7 of 10

    Board approved policy, a risk register entry, quarterly reporting. Day to day overseers are unnamed and their authority to halt is assumed.

  6. 06

    AI Integration

    5 of 10

    Writes to the system of record with an audit trail, under a service account. The originating caller is lost at the boundary.

  7. 07

    Autonomy Envelope

    6 of 10

    A written boundary, a value threshold and a hard stop. The stop has never been exercised and there is no rollback path.

Why the agent is invented.

A methodology never applied in public is a claim rather than an instrument. This page closes that gap in the only way that can be done honestly today.

A reader can study the seven dimensions and the five tiers and still not know what a score of 58 means, what evidence produces a 4 rather than a 7, or why two deployments with the same total can land in different tiers. Scoring a real, named AI product would have been easy, and the exercise would have been invalid.

Read what the dimensions ask for: a documented blast radius per tool, a named individual with written authority to halt the agent, a retained run history with dates and versions, per caller spend caps, an audit trail that keeps the originating identity across a service boundary. None of that can be observed from outside for any third party system. A vendor's public documentation describes what a product can do, never what a deployer's evidence file contains. A score built on absent evidence would measure how much a company chooses to publish, dressed in the apparatus of a real assessment.

It would also be a reputational claim about a company that never asked for one, made by a party with no relationship to it, and the words "illustrative only" do not survive being quoted by a search engine or an assistant three steps downstream. So the deployment is constructed. Nothing empirical is asserted about anyone, and every number comes from the published rubric, which means the arithmetic can be checked instead of trusted.

The deployment being scored.

A claims triage agent at a mid sized European insurer. Everything about it is invented.

The agent reads a first notification of loss, retrieves the relevant policy wording and prior claim history from an internal document store, classifies the claim and drafts a coverage position. It then routes the file to a handler or, below a stated value threshold, sends the customer a standard acknowledgement itself.

It calls six tools: the policy document store, the claims system of record, an email service, a fraud screening API, a payments read endpoint and an internal knowledge base. It has been in production for eleven months. Each strength and each gap below was chosen because the rubric asks for that specific control, and no particular insurer looks like this.

The arithmetic, line by line.

Each dimension is scored from 0 to 10 against the rubric in Methodology v2.0. The sub score times the weight, divided by ten, is its contribution. The seven contributions add up to a total out of 100.

Constructed example. No real organisation is described in this table.
DimensionWeightSub scoreContribution
Trust and Safety186 of 1010.8
Context Integrity145 of 107.0
Distribution Control124 of 104.8
Product Maturity147 of 109.8
Governance167 of 1011.2
AI Integration125 of 106.0
Autonomy Envelope146 of 108.4
Weighted total10058

Constructed example. The weights are fixed and published: Trust and Safety 18, Governance 16, Context Integrity 14, Product Maturity 14, Autonomy Envelope 14, Distribution Control 12, AI Integration 12. A floor applies at every tier, set out on certification levels: at least 4 on every dimension for Certified, 6 for Advanced, 8 for Elite. The floor is the part most readers miss, and it decides this example.

What produced each number.

For every dimension, the evidence that was produced and the evidence that was not. The second list moves the score, because an assessment scores what can be shown.

Trust and Safety 6 of 10, weight 18

An external firm ran a red team exercise before launch and repeated it once since. Refusal behaviour and misuse detection are implemented and tested, and findings from the second exercise are logged with owners.

What it costs. Both exercises used the same scenario set, so the second measured regression and found no new attack surface. Production misuse is reviewed by sampling, without continuous monitoring.

Context Integrity 5 of 10, weight 14

The retrieval corpus is a controlled internal document store with a named owner and a documented refresh job.

What it costs. Freshness is asserted by the refresh job and never measured. No staleness signal reaches the agent or the reviewer, and no provenance record travels with a retrieved passage, so a reviewer reading an output cannot tell which version of a document produced it.

Distribution Control 4 of 10, weight 12

Invocation is authenticated per caller, with permissions managed through the organisation's single sign on. Development and production are separated with distinct credentials.

What it costs. There are no per caller rate limits or spend caps, and blast radius is documented for two of the six tools. On the widest of the six, nobody has established the ceiling on records affected in a single run.

Product Maturity 7 of 10, weight 14

Versioned releases, a regression suite on every prompt and model change, uptime measured against a stated target, and run history kept for eighteen months.

What it costs. The held out task set comes from the same pool as the development examples, so the suite measures consistency and says little about generalisation.

Governance 7 of 10, weight 16

A written AI policy approved by the board, an entry in the enterprise risk register with a named risk owner, quarterly reporting to a risk committee, and an incident protocol rehearsed once.

What it costs. Accountability stops at the risk owner. The people performing day to day oversight are named in no document, and their authority to halt the agent is assumed, never granted in writing.

AI Integration 5 of 10, weight 12

The agent writes to the system of record through a service account with an audit trail, and escalations land in the existing case management queue.

What it costs. The audit trail records the service account, so the originating caller is lost at the boundary. An escalation arrives without the reasoning or the retrieved context behind it, and the handler who picks it up starts from zero.

Autonomy Envelope 6 of 10, weight 14

The boundary between autonomous action and human confirmation is written down, a value threshold triggers confirmation, and a named team can trigger a documented hard stop.

What it costs. The hard stop has never been exercised outside a design review, and there is no rollback path for actions already written to the system of record. The envelope is specified and untested.

Why 58 stops at Certified.

On the number alone the deployment sits inside the Advanced band, which runs from 55 to 74. It does not reach Advanced.

58Weighted total, inside the Advanced band.
3Dimensions below the Advanced floor of 6: Distribution Control at 4, Context Integrity and AI Integration at 5.
CertifiedThe tier awarded. Its floor is 4, which Distribution Control meets exactly.

This is the point of the floor, and the reason the methodology is no simple average. The three dimensions that fail are the three that govern what happens when something goes wrong: who could invoke the agent and how far one access event reaches, whether the agent can tell current information from stale, and whether a person picking up an escalation can see why it was raised.

The four dimensions that score well are the ones an organisation builds first, because a board asks about them. An average would let good governance paperwork carry a weak access model into a tier it has not earned. That is the ordinary shape of a competent organisation that started from policy and has not yet reached the plumbing.

For the operator, the fastest route from Certified to Advanced here runs through engineering, ahead of any further governance work. Per caller rate limits. Blast radius documentation for the remaining four tools. A provenance record that travels with retrieved passages. The originating identity kept across the service account boundary.

Four pieces of engineering, none of them large, worth more than another policy document.

Has any organisation been certified under this methodology?

As of 17 August 2026, when this page was published, no organisation had been assessed or certified under Agent Certified, and no score had been published for any real party.

The methodology, the weights, the rubric and the tier rules are published in full so they can be read and argued with before anyone is scored against them, which is the order in which we think a standard should be built. A first published assessment will carry the assessed party's agreement, describe the evidence file, and say so plainly on this page.

How should an operator use this page?

Go through the seven dimensions above and ask, for each, which side your own deployment sits on.

The evidence produced is deliberately unremarkable, because most of it already exists in a competent organisation and has simply not been gathered in one place. The evidence missing is where an assessment spends its time. An operator who works through those seven gaps before commissioning anything will see a materially different first score, and the exercise costs nothing but attention.

The full rubric is in Methodology v2.0, the tier rules are on certification levels, and the sequence of an assessment is in the assessment process guide.

Your own agent

Find out its shape from the evidence.

One named agent, scored against the same rubric, with a tier and a report your insurer, board and buyers can read. Four to six weeks from intake.