- The prompt governs what an agent is told to do. The retrieval corpus governs what it is told is true. Only the first is routinely under change control, and the second is where wrong answers come from.
- Overbroad retrieval scope is the most common finding. Asked what the agent can read, teams describe an intention. Asked to enumerate it, they find draft pricing, old contracts and a superseded handbook in the same folder tree.
- Without a documented authority hierarchy, two contradictory documents both stay retrievable and the phrasing of a question decides which one the customer is told. That is the shape of the failure in Moffatt v. Air Canada.
- Three corpus events are behavioural changes and are almost never logged: a document added, a document edited in place, and a document deleted. The middle one is the sharpest, because the index updates silently.
- Re-embedding changes behaviour with every document byte-identical. Chunk size, overlap and the embedding model are system parameters, not infrastructure details.
- The corpus falls between the frameworks. Article 10, ISO/IEC 42001 and the NIST AI RMF are oriented at the data a model was built from. The corpus is the data a deployment reasons over, and it is the deployer's alone.
Ask an operator how their agent knows anything and you get one of two answers. Either the model knows it, which is a description of a system that will confidently state things about your business that are not true, or the agent looks it up, which is the right architecture and the beginning of a governance problem nobody has named.
Retrieval solved the most visible failure mode in deployed language systems. An agent grounded in your own documents stops inventing your refund policy and starts quoting it. What that architecture does, and what almost no organisation has noticed it doing, is take the accuracy of the system and make it a property of a document repository maintained by people who do not know they are maintaining a system component.
This article sets out what an assessment looks for in that repository. It is written for operators preparing an evidence pack, and for the procurement and underwriting readers who will be handed the result. It is a companion to prompt change control, which covers the instruction layer, and it belongs alongside it: an organisation with excellent change control over its prompts and none over its corpus has half a system under management.
Five properties an assessment scores
The corpus is examined as a controlled component, in the same register as any other. Five properties carry the weight.
1. Inventory and boundary
The first question is the simplest and the one that most often stops the conversation: produce the list of everything this agent can read.
The intended answer describes categories. The policy documents. The product pages. The support knowledge base. The actual answer, when someone goes and looks, is usually a path. The retrieval index was pointed at a directory, or at a workspace, or at a document management system with a permission model that predates the agent by a decade. The directory has subdirectories. Somebody put a pricing draft in one of them in March.
This is not carelessness in the ordinary sense. It is the consequence of the corpus never having been recognised as a permission surface. A folder is a place to keep documents. A retrieval scope is a statement about what a system is entitled to assert. The two were the same object and only one of them had a threat model.
What scores well is a boundary that is stated positively rather than negatively: an enumerated set of sources with named owners, rather than a location with an assumption that nothing inappropriate will be added to it.
2. Authority and precedence
The second property is the one with the case law attached.
Real organisations hold contradictory documents. The published policy page says one thing, the internal guidance note says something slightly different, last year's handbook says a third thing and has not been withdrawn because nobody withdraws handbooks. Human staff resolve this constantly and invisibly, by knowing which source is the real one and which is a document somebody wrote once.
A retrieval system has no such knowledge. Both documents are text, both are retrievable, and which one surfaces depends on how the question was phrased. The customer gets whichever answer the index preferred that morning.
This is the shape of the failure in Moffatt v. Air Canada, decided by the British Columbia Civil Resolution Tribunal in 2024, where a chatbot gave a customer information about a bereavement fare that conflicted with the airline's own published policy page, and the airline was held responsible for what its chatbot had said. The instructive part for governance is not the outcome, which was unsurprising. It is that two authoritative statements coexisted inside one organisation and nothing in the system decided which was operative.
An assessment therefore asks for an authority hierarchy: which source is canonical for which class of question, what rank each source holds, and what the system does when a lower-ranked source contradicts a higher-ranked one. Most organisations do not have this written down anywhere, and writing it down is frequently the single most useful thing an assessment causes to happen.
3. Currency and expiry
Documents go out of date. Retrieval systems do not notice.
A price list superseded in January remains a perfectly well formed, highly retrievable document in August. A policy amended after a regulatory change sits next to its own previous version, both plausible, both indexed. In a human process the old version is stale and looks stale, because the person reading it recognises the letterhead or the date or simply the fact that this is not the document everyone uses. In a retrieval process it is a passage of text with good semantic overlap with the question, which is the only property that matters.
Two controls score here, and they are different. The first is an expiry or review interval per source, so that currency is a property of the document rather than an assumption. The second is a withdrawal process, so that superseded material leaves the retrieval scope rather than merely being superseded in someone's understanding. The second is rarer than the first and matters more.
4. Provenance
For each source: where it came from, which system of record it belongs to, who owns it by name or role, and when a person last reviewed it. This is unglamorous and it is what makes the other four properties auditable rather than asserted.
It also does specific work at incident time. When an agent has said something wrong, the first question is whether it said something the corpus contained or something it produced. Those have entirely different remediations, and answering the question at all requires knowing what the corpus contained on the day. An organisation that cannot reconstruct its corpus as at a past date cannot investigate its own incidents, which is a finding under post-market monitoring before it is anything else.
5. Exclusion
The fifth property is the inverse of the first, and it is scored separately because the failure it prevents is a different failure. Not what should the agent be able to read, but what must it never be able to read.
Personnel records. Draft commercial terms. Legal advice. Unredacted contracts with third parties. Board material. Anything under a confidentiality obligation to somebody else. In a retrieval architecture these are not protected by being sensitive. They are protected only by being outside the index, which is a configuration state that somebody can change in an afternoon with the best of intentions, usually to fix a case where the agent could not answer a question.
What scores is an explicit exclusion list with a change process attached, so that widening scope is a decision with a name on it rather than an incremental convenience.
The three events nobody logs
Change control over the corpus fails differently from change control over a prompt, because the changes do not look like deployments. Three events are behavioural changes to the system and are almost never recorded as such.
A document is added. The most visible of the three, and the one most likely to be noticed, because somebody deliberately put something in. It is still rarely logged against the agent, because the person adding it is publishing a document, not modifying an AI system, and in their own frame of reference they are right.
A document is edited in place. This is the sharpest of the three. Nobody added anything, nobody removed anything, and the agent's answer to a recurring question changed at whatever hour the index refreshed. There is no deployment, no release note, and no event anywhere in the engineering record. An operator investigating a behaviour change will examine the prompt, the model version and the parameters, find all three unchanged, and reasonably conclude that the provider changed something. The provider did not. A colleague corrected a sentence.
A document is deleted or falls out of scope. The failure here is silence. An agent that can no longer retrieve the authoritative answer does not say so. It answers from whatever remains, which is by definition a lower-authority source, with exactly the same confidence it had the day before. Deletion is the corpus event with the worst ratio of visibility to consequence.
The change that touches no document
There is a fourth event, and it is the one that most often defeats an otherwise well governed deployment, because it changes behaviour with every document in the corpus byte-identical before and after.
Retrieval does not read documents. It reads fragments of them, produced by splitting each document into chunks and converting each chunk into a vector using an embedding model. Which fragments surface for a given question is determined by how the splitting was done and by which embedding model did the conversion.
Change the chunk size and a policy clause that previously arrived with its qualifying sentence now arrives without it. Change the overlap and a definition separates from the term it defines. Change the embedding model, or accept a new version of it, and the entire similarity geometry of the corpus shifts, which alters what surfaces for questions nobody thought to re-test.
None of this is exotic and all of it is routinely treated as infrastructure maintenance rather than as a change to a system that makes consequential statements. An assessment treats chunking configuration and embedding model version as system parameters, subject to the same recording discipline as the system instruction, with a regression run after any change. The methodology for testing behaviour that does not repeat itself is set out in auditing dynamic behaviour.
Why the frameworks miss this
Operators who have done serious governance work sometimes push back at this point, reasonably: we have addressed data governance, it is in our evidence pack, we mapped it to the standards. Almost always what has been addressed is training data.
That orientation runs through the instruments. The EU AI Act's data governance provisions at Article 10 are addressed to training, validation and testing data sets for high-risk systems, and the deployer implications are set out on agentliability.eu. ISO/IEC 42001 and the NIST AI Risk Management Framework both treat data as an input to system development. The questions they ask are good questions and they are asked about the wrong artefact for this purpose, because for a deployer using a general-purpose model the training data decisions were somebody else's, taken years ago, and are largely unobservable.
The retrieval corpus is the opposite in every respect. It is the deployer's own. It changes weekly. It is edited by people with no engineering role. And it determines what the agent asserts as fact this afternoon. A framework crosswalk that covers training data governance thoroughly and says nothing about the corpus has examined a decision taken elsewhere and left the live component unexamined.
This is not a criticism of the standards, which were written when the dominant architecture was a trained model answering from its own weights. It is an observation about where the gap currently sits, and the practical consequence is that an operator who wants corpus governance has to build the evidence deliberately, because no questionnaire will ask for it. The wider question of what transfers between frameworks is treated at agentliability.co, on reusable governance evidence.
The minimum viable corpus register
Seven fields per source, held next to the change record for the prompt and the model. For a mid-sized deployment this is a spreadsheet and an afternoon, and it is the artefact that carries the most weight per hour spent.
- What the source is, at the level of a document or a defined collection.
- Which system of record it comes from, and by what mechanism it reaches the index.
- Who owns it, by name or role, meaning who is accountable for its content being right.
- When it was last reviewed, and by whom.
- Its authority rank, and for which classes of question it is canonical.
- Its expiry date or review interval, and what happens when that date passes.
- Whether it is inside or outside retrieval scope, with the date and authoriser of the last scope change.
Add to that a versioned snapshot of the corpus taken on a defined interval, so that the state of the knowledge base on any past date can be reconstructed. That single control is what converts an incident from an argument into an investigation.
How it is scored
Corpus governance is read primarily into Data Governance, with the authority hierarchy and the currency controls also read into Trust and Transparency, because an agent asserting a superseded policy is a transparency failure as much as a data one. Regression evidence following a chunking or embedding change is read into Performance and Reliability. Exclusion scope sits with Security and Resilience.
It is not a dimension of its own, for the same reason change control is not. It is a mechanism that determines whether statements made in other dimensions remain true. A strong finding on transparency means very little if the document the agent is transparently quoting was withdrawn in March.
Two audiences use the result. The first is a procurement or audit function that wants to know whether an assertion about the agent has a shelf life. The second is an underwriter, for whom a corpus register and a snapshot history do the same work a change record does: they convert an opaque operating period into something that can be examined after a loss. The connection between evidence quality and terms is set out at how certification feeds underwriting, and the recovery consequences of poor evidence at agentinsured.eu, on subrogation and AI vendor contracts.
The finding that surprises people
Operators expect the corpus conversation to be about size, and arrive prepared to explain that their knowledge base is large and comprehensive. Size is close to irrelevant to the assessment. A tightly scoped corpus of forty owned, dated, ranked documents scores substantially better than forty thousand documents in a folder nobody has audited, and it produces a better agent.
The reason is that a retrieval system's failure mode is not missing information. It is confidently surfacing the wrong information, and every additional unowned document is another opportunity for that. Comprehensiveness without governance is not a strength being under-recognised. It is the risk itself, described approvingly.
Questions
Why does a certification assessment look at the retrieval corpus separately from training data?
Because they are governed by different people on different timescales and only one is covered by the existing frameworks. Training data governance asks how a model was built, is answered once, and for a deployer using a general-purpose model it is somebody else's answer. The retrieval corpus is the deployer's own document set, it changes weekly, it is edited by people with no engineering role, and it determines what the agent asserts as fact today. An assessment that reviewed training data governance and stopped has examined a decision taken elsewhere and ignored the component that will produce next month's incident.
What is the most common finding on retrieval scope?
That nobody can produce the list. Asked what the agent can read, teams describe an intention: the policy documents, the product pages, the support articles. Asked to enumerate it, they find the index was pointed at a folder, the folder has subfolders, and the subfolders contain draft pricing, an old contract and a superseded handbook. Overbroad scope is not a configuration error in the usual sense. It is the absence of anyone having treated the corpus as a permission surface, which is what it is.
What happens when two documents in the corpus contradict each other?
Without a precedence rule, whichever one the retrieval step surfaces wins, and which one that is can change with a question's phrasing. This is the shape of the failure in Moffatt v. Air Canada, where a chatbot gave a customer information about a bereavement fare that conflicted with the airline's own published policy page and the airline was held responsible for what the chatbot said. Two authoritative statements coexisted and nothing decided which was operative. An assessment asks for a documented authority hierarchy: which source is canonical for which class of question, and what happens when a lower source contradicts a higher one.
Does re-embedding a corpus count as a change to the system?
Yes, and it is the change most often missed because no document was touched. Retrieval depends on how documents were split into chunks and on which embedding model turned those chunks into vectors. Changing chunk size, overlap, the embedding model or its version alters which passages surface for a given question, and therefore what the agent says, with every document byte-identical before and after. It should be recorded with the same discipline as a change to the system instruction, and a regression run should follow it.
What does a minimum viable corpus register contain?
Seven fields per source. What the document or collection is. Which system of record it comes from. Who owns it, by name or role. When it was last reviewed and by whom. Its authority rank for the classes of question it answers. Its expiry or review interval. And whether it is in or out of retrieval scope, with the date of the last scope change. Held next to the change record for the prompt and the model, this is what allows an assessment statement to remain true after the assessment.
Is a larger knowledge base better for an assessment?
No, and the assumption is worth correcting because acting on it makes the result worse. A tightly scoped corpus of forty owned, dated, ranked documents scores better than forty thousand documents in an unaudited folder, and produces a better agent. A retrieval system's failure mode is not missing information but confidently surfacing wrong information, and every unowned document is another opportunity for that. Comprehensiveness without governance is the risk, described approvingly.
Sources
- European Parliament and Council. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence, Articles 10, 12, 13, 14 and 72. OJ L, 12.7.2024. Article 10 addresses data governance for training, validation and testing data sets used by high-risk AI systems.
- Regulation (EU) 2026/1744, the AI Omnibus. Entered into force 27 July 2026. Annex III obligations from 2 December 2027, Annex I from 2 August 2028. European Commission, AI Omnibus enters into force.
- Moffatt v. Air Canada, British Columbia Civil Resolution Tribunal, 2024. Air Canada held responsible for information its chatbot provided to a customer which conflicted with the airline's own published policy page.
- International Organization for Standardization. ISO/IEC 42001:2023, Information technology, Artificial intelligence, Management system. Geneva, 2023. No clause identifiers are cited here; the normative text is a commercial publication and is not quoted.
- National Institute of Standards and Technology. AI Risk Management Framework 1.0 (NIST AI 100-1), released 26 January 2023. nist.gov.
- Agent Certified. Methodology specification, published at agentcertified.eu/methodology. Source for the dimensions referenced above and for the evidence tiers applied in assessment.