Key takeaways
- A low certification score places a system in the Pre-Assessment band, below 20 out of 100, and produces a private findings document, not a public failing grade or a notice sent to any third party.
- Most first-time assessments do not reach a passing level. Undocumented or informally enforced governance, not technical capability, is the most common reason.
- The three dimensions that most frequently produce low-scoring findings are Governance, Autonomy Envelope, and Trust and Safety, each of which depends on documentation and enforced controls rather than raw model quality.
- Re-assessment after remediation is scoped to the dimensions that produced findings, not repeated from a blank state, which typically makes the second pass faster and less resource-intensive than the first.
- Nothing about a low score prevents an operator from later achieving a higher level once the specific, named gaps are closed. The findings document is a work plan, not a verdict.
The question behind the question
Operators considering an assessment often hesitate for a reason they rarely say out loud: what if the agent fails. It is a reasonable thing to worry about, and it deserves a direct answer rather than reassurance. Failing, in the sense of scoring below a passing threshold on a first assessment, is common. It is not a special or embarrassing outcome. The framework assumes that an AI agent deployment will not arrive with governance documentation, tested behaviour boundaries and an enforced autonomy scope already in place, because until recently almost nobody has been asked to produce that evidence. An assessment is often the first time anyone has systematically asked the questions a certification requires answers to. A low score usually reflects that gap, not a defect in the underlying AI system.
What a low score actually is: a findings document, not a rejection
The output of every assessment under this framework is the same regardless of the score: a scorecard across the seven dimensions, a written breakdown of the evidence reviewed for each, and a findings document identifying specific gaps. Where the combined score falls below 20, the resulting certification level is Pre-Assessment, the entry tier in this framework's five-level structure (Elite, Advanced, Certified, In Progress, Pre-Assessment). Pre-Assessment is not a rejection code. It is a level like any other in the structure, and the findings document that accompanies it is written the same way a findings document at any other level is written: specific, evidenced, and organised by dimension so the operator knows exactly what changed the outcome.
This matters because operators sometimes assume, reasonably given how certification language works in other industries, that a failed assessment is recorded somewhere as a permanent mark against the system or the organisation. It is not. There is no failing designation distinct from a low score, no public register of assessments that did not pass, and no requirement to disclose a Pre-Assessment result to anyone. The assessment record belongs to the operator who requested it.
Where the record goes, and where it does not
The scorecard, dimension breakdown, and findings document are delivered directly to the operator that requested the assessment. They are not published, not shared with the public-facing side of this site, and not disclosed to any third party, including an insurer, a regulator, or a prospective client, unless the operator chooses to share them. This is the same confidentiality standard applied to every assessment regardless of outcome; a high-scoring assessment is held to the identical default of operator ownership until the operator elects to publish it, most commonly because the resulting level supports a claim the operator wants to make publicly.
The practical consequence is that an operator can request an assessment, receive a low score, and treat the exercise entirely as internal diagnostic work with no external visibility at all, right up until they choose to request a re-assessment and, eventually, decide whether to make a passing result public. This is a deliberate design choice, consistent with the position this framework has taken elsewhere on being honest about what evidence exists and does not: see the companion analysis of whether certification actually reduces insurance premiums, which applies the same principle of separating what is demonstrated from what is merely likely.
What actually causes a low score
Three dimensions of the seven-dimension framework account for the large majority of findings that keep a first assessment in the Pre-Assessment or In Progress band.
Governance. The most common single gap is the absence of a named, accountable owner for the AI system's outcomes, combined with no documented incident response process and no record of who authorised the deployment in the first place. Our earlier analysis of the Governance dimension sets out the evidence standard in detail; in practice, most first-time submissions can describe governance informally in conversation but cannot produce the written record an assessor needs to score it.
Autonomy Envelope. The second most common gap is a scope of autonomous action that was never precisely written down, or that exists in a policy document but is not actually enforced by the system itself. An agent that is described as being limited to a certain category of action, but whose underlying configuration does not technically prevent it from doing more, scores poorly here regardless of how the system has behaved so far in practice, because the assessment evaluates what the system is capable of doing, not only what it has done to date.
Trust and Safety. The third common gap is behaviour boundaries that exist as an informal understanding within the team but have not been tested against adversarial or edge-case inputs, and are not documented in a form that survives staff turnover or an external audit. Systems that behave well in normal use frequently have not been stress-tested against the inputs an assessor, or eventually an underwriter, will ask about directly.
Technical model quality is rarely the primary driver of a low score. Two organisations running functionally similar underlying models can score very differently depending entirely on whether governance, scope, and safety testing were documented before the assessment began.
The remediation and re-assessment path
Remediation is scoped directly to the findings document, not to the assessment as a whole. Findings that are primarily documentation gaps, writing down an approval process that already happens informally, or producing a written incident response plan from an existing but unrecorded practice, are typically the fastest to close, often within two to four weeks. Findings that require building genuinely new controls, implementing an enforced approval threshold in the system's configuration rather than describing one in a policy, or running a structured adversarial test suite for the first time, typically take six to twelve weeks depending on engineering capacity.
Re-assessment following remediation does not start from a blank state. It is scoped to the dimensions that produced findings in the original assessment, provided the underlying system has not changed materially since. Dimensions that scored adequately the first time, commonly Context Integrity, Product Maturity, and Distribution Control for well-built systems with weak governance documentation, are carried forward rather than re-evaluated from scratch. This is why a second assessment following genuine remediation work is typically both faster and less resource-intensive than the first, even where multiple dimensions required real changes rather than paperwork alone.
Where an operator's system involves multiple agents calling each other, remediation can also surface additional scope that was not visible in the original submission; see the companion analysis of certifying multi-agent systems for how findings can expand once a fuller call graph is produced during remediation.
What to tell a board or an underwriter about a low first score
The honest and, in this framework's experience, commercially unremarkable position to take is that a low first-time score is the ordinary starting point for most organisations undertaking this kind of assessment for the first time, and that what matters is the trajectory from findings to remediation to re-assessment, not the initial number. An underwriter or a board member who understands how certification assessments generally work does not expect a first attempt to be a passing one; what they are actually evaluating, whether they say so explicitly or not, is whether the organisation treats the findings as a genuine work plan or ignores them. Operators preparing for an assessment for the first time, and wanting to understand what evidence to gather before starting rather than after receiving findings, should read preparing for an AI agent certification assessment. For how a completed, passing assessment then feeds an insurance underwriting conversation, see preparing an AI agent underwriting submission on agentinsured.eu.
Frequently asked questions
What happens if my AI agent fails certification?
It does not become a public rejection. A low-scoring assessment produces a findings document that identifies exactly which of the seven dimensions fell short and why, alongside a scorecard that places the system in the Pre-Assessment band, below 20 out of 100. That document is the operator's property. It is not published, and no failure notice is issued to any third party. The realistic path forward is remediation against the specific findings, followed by re-assessment once the gaps are closed, not a one-time pass or fail verdict.
Is a low certification score made public or shared with anyone?
No. The scorecard, dimension breakdown, and findings document produced by an assessment are delivered to the operator that requested it and are not published or disclosed to any third party, including insurers, regulators, or the public listing, unless the operator chooses to share them. Only assessments the operator elects to make public, typically because the score supports a certification level the operator wants to display, appear anywhere outside the operator's own records.
What are the most common reasons an AI agent scores low on certification?
The most frequent gaps cluster in three dimensions: Governance, where systems lack a named accountable owner, a documented incident response process, or a record of who approved the deployment; Autonomy Envelope, where the scope of autonomous action the agent is authorised to take was never written down or is not enforced in the system itself; and Trust and Safety, where behaviour boundaries exist informally but have not been tested or documented in a form an assessor or an underwriter can verify. Technical capability is rarely the primary cause of a low score. Undocumented governance is.
How long does remediation take after a failed assessment?
It depends entirely on which dimensions produced the low score. Findings concentrated in documentation gaps, such as writing down an existing but undocumented approval process, can often be remediated in two to four weeks. Findings that require building new controls, such as implementing enforced approval thresholds in the system itself or establishing a genuine incident response process where none existed, typically take six to twelve weeks. The findings document produced at assessment specifies which category each gap falls into, so operators can estimate their own remediation timeline rather than guessing.
Can I request a partial or re-assessment instead of starting over?
Yes. A re-assessment following remediation is scoped to the dimensions that produced findings in the original assessment, not repeated in full from a blank state, provided the underlying system has not changed materially since the original assessment. Dimensions that scored adequately the first time are carried forward rather than re-evaluated from scratch, which is why remediation and re-assessment is typically faster and less resource-intensive than the initial assessment, even where several dimensions required rework.
References
- Agent Certified. Methodology specification and certification levels (Elite, Advanced, Certified, In Progress, Pre-Assessment), published at agentcertified.eu/methodology and agentcertified.eu/certification-levels.
- Agent Certified. Seven-dimension framework: Trust and Safety, Context Integrity, Distribution Control, Product Maturity, Governance, AI Integration, Autonomy Envelope.
- International Organization for Standardization. ISO/IEC 42001:2023, Information technology, Artificial intelligence, Management system, referenced as a comparable governance evidence standard.
For how an insurer reads a certification record once remediation is complete, see AI agent certification and insurance eligibility on agentinsured.eu, and for the SME-scale version of building the same evidence from scratch, see the pre-deployment insurance checklist on insureyouragent.com.