In short
  • For an agent built on a general-purpose model, the prompt is most of the system. Certifying behaviour that cannot be pinned to a version certifies a snapshot with no shelf life.
  • Seven surfaces change an agent's behaviour: system instruction, model version, inference parameters, tool permissions, retrieval sources, guardrail configuration and escalation thresholds. Most organisations track one of the seven, and it is usually not the prompt.
  • Frequency of change is not penalised. Unrecorded change is. Forty logged edits score better than two unlogged ones.
  • Provider-side model updates change behaviour without any action by the operator. Two answers score: version pinning with a documented adoption process, or a scheduled regression suite that would detect the drift.
  • A minimum viable change record has six fields per entry and takes minutes. It is the artefact that most improves both an assessment result and an insurance submission, because it converts an opaque operating history into a timeline.
  • Change control is scored inside Governance and the Autonomy Envelope, with regression evidence read into Performance and Reliability. It is not a dimension of its own because it is what makes every other dimension verifiable over time.

An assessment team spends most of its time establishing what an agent does. The harder question, and the one that decides whether the answer is worth anything, is how long the answer stays true.

A conventional software audit has a straightforward answer to this. The system is a build, the build has a hash, and the statement attaches to the hash. An AI agent built on a general-purpose model does not work that way. The application code is often the least behaviourally significant part of it. What the agent will say, what it will refuse, what it will do without asking a human, and how it handles the cases nobody anticipated are set by a body of natural language instruction, a set of permissions, a model identifier and a handful of numeric parameters. All of those can be changed in a text field by someone with no engineering role, and in most organisations they regularly are.

That is the practical subject of this article. It is written for operators preparing for assessment, and for the procurement, audit and underwriting readers who will be handed the result and need to know what it is warranting.

The seven surfaces

Before a change record can be assessed, the organisation has to agree on what counts as a change. In our experience this conversation alone surfaces more than the document review that follows, because teams routinely discover that three of these surfaces belong to nobody.

Surface Why a change here changes the agent
1. System instruction and prompt templates The primary determinant of behaviour, tone, refusals and scope. Usually the most frequently edited surface and the least often versioned.
2. Model identifier and version A different model is a different system. This surface can change without any action by the operator, which is why it needs its own control.
3. Inference parameters Sampling settings alter the variance of output. A change here does not change what the agent can say, it changes how often it says the unusual thing.
4. Tool and action permissions The boundary of the Autonomy Envelope. Adding a write action to an agent that previously only read is the single most consequential change type in this list.
5. Retrieval sources An agent grounded on a document set behaves according to that set. Adding, removing or letting a source go stale changes answers without touching a line of instruction.
6. Guardrail and filter configuration Frequently loosened in response to false refusals, frequently loosened by someone under time pressure, and almost never recorded as a risk decision.
7. Escalation and human review thresholds Determines when a person sees the output. Raising a threshold to clear a backlog silently converts a supervised process into an unsupervised one.

Surfaces four, six and seven are where an assessment finds the most consequential unrecorded changes, and they share a pattern. Each of them is usually adjusted for an operational reason that is entirely reasonable at the time, by someone who is not thinking about it as a change to a governed system. The permission was added to close a ticket. The filter was relaxed because it kept blocking a legitimate phrase. The review threshold was raised because a queue was backing up on a Friday. None of those are bad decisions. All of them are risk decisions, and a governance framework that cannot see them is not governing the agent.

The test we apply. Ask the operator to state, in one sentence, what their highest-consequence agent was permitted to do without human review on a specific date three months ago. Then ask them to prove it. The answer to the first question is usually confident. The answer to the second is where the assessment actually begins.

Why the regulation presupposes this and does not say it

No provision of the EU AI Act uses the phrase prompt change control. Several provisions are unworkable without it.

Record-keeping under Article 12 exists so that the operation of a system can be reconstructed after the fact. A log of what the agent output, without a record of what the agent was instructed to do at that moment, reconstructs half of an event. Human oversight under Article 14 requires that a natural person be able to oversee the system effectively, which presumes the person knows what the system currently is. Post-market monitoring under Article 72 asks an organisation to observe its system's behaviour in the field and act on what it observes, and observing behaviour that changes for unrecorded reasons produces a monitoring record that cannot support a conclusion. Deployer duties under Article 26 rest throughout on the deployer knowing the configuration it is operating.

The timing has shifted but not the substance. The Digital Omnibus, in force since 27 July 2026, moved the Annex III high-risk package, which carries the record-keeping and post-market monitoring duties, to 2 December 2027, and the Annex I package to 2 August 2028. Article 50 transparency obligations were not deferred and applied from 2 August 2026. A deployer with more runway to build these processes has more runway, not a smaller obligation. Our reading of what survived the deferral is at whether the delay means you can skip certification, and the operator-level analysis is on agentliability.eu.

The management system standards are more direct. ISO/IEC 42001:2023 is an AI management system standard, and change management is a structural expectation of any management system: an organisation that cannot demonstrate control over changes to the thing being managed does not have a management system, it has a description of one. The NIST AI Risk Management Framework, released in January 2023, places this in its MANAGE function, where risks that have been identified and measured are acted on and monitored over time. Monitoring over time is only meaningful against a known baseline. Our crosswalk between these frameworks is at the ISO 42001 and NIST AI RMF control mapping.

The provider-side change nobody authorised

The surface that most often defeats an otherwise good change process is the one the operator does not control. A model provider updates a model. Nothing in the operator's estate changed. The agent behaves differently.

An assessment does not treat this as a fault. It treats it as a control question with two acceptable answers.

The first is pinning. Where the provider offers a stable version identifier, pin to it, and treat adoption of a new version as a change requiring evaluation, testing and a dated record, rather than something that happens to you. This converts an external event into an internal decision.

The second, where pinning is not available or where an organisation has accepted a rolling version, is detection. A regression suite of representative inputs with expected behavioural properties, run on a schedule, with results retained. This does not prevent the change. It ensures the organisation finds out, and it produces the dated evidence that shows when behaviour shifted.

An organisation that can offer neither is in a specific position that should be stated plainly rather than scored around: it cannot say what its agent does today, only what it did when someone last looked. That is not a failing grade by itself, and it does place a ceiling on what any certification statement about that agent can responsibly claim. The dependency question more broadly is treated in certifying agents built on foundation models.

What the evidence looks like, and how it scores

The assessment asks for the change record covering the period under review across all seven surfaces. What comes back sorts into four broad tiers.

Scores nothing. A repository history for the application that does not contain the prompt. A shared document of prompt variants with no dates and no indication which is live. Screenshots. A recollection, however confident, of when a permission was added.

Scores partially. A dated change list that records what changed but not who authorised it or what was tested. This establishes a timeline, which is genuinely useful and is more than most organisations have. It does not establish that anyone weighed the change before it reached production.

Scores well. Prompts and configuration held in version control alongside the application, so that every change carries an author, a timestamp and a diff by construction rather than by discipline. Paired with a record of what test was run before release and who approved it. The key property here is that the evidence is a by-product of the working method rather than a document produced for the assessment, which is also why it survives contact with a busy quarter.

Scores at the top of the range. The above, plus a demonstrated rollback. An organisation that can show a date on which a change was made, detected as harmful, and reverted, with the detection method and the elapsed time recorded, has evidenced the entire control loop rather than one half of it. This is rare and it is worth more in an assessment than a longer period of uneventful operation.

The minimum viable record

Small teams reasonably object that a formal change process is disproportionate to a six-person company running one support agent. It is. What follows is not a process, it is six fields, and it takes minutes per change.

  1. When. Date and time the change reached production, not the date it was written.
  2. What surface. One of the seven above.
  3. From and to. The previous and new value, or a reference to a stored diff.
  4. Who authorised it. A person, not a team.
  5. What was tested. Even if the honest answer is a named set of five manual checks.
  6. How it reverts. The rollback path, stated before it is needed.

Kept in the same version control as the application, this satisfies the substance of what an assessment asks for at any organisation size. Kept in a spreadsheet, it still satisfies most of it, and is considerably better than the alternative. The failure mode to avoid is a separate governance document maintained by someone who is not making the changes, which decays within a quarter and then actively misleads.

What this is worth outside an assessment

Two audiences beyond the assessor read this artefact, and both read it closely.

The first is the incident responder, who is usually the same person on a much worse day. The first useful question after an AI failure is what changed and when, and an organisation that can answer it in minutes contains the incident inside a working day. An organisation that cannot spends that day reconstructing its own configuration history from memory and chat logs, while the behaviour continues.

The second is the underwriter. Insurance for AI exposures is written on a claims-made basis, which means the temporal reach of the policy is negotiated rather than assumed, and the operating period before inception is the period the underwriter cannot see. A change history is the artefact that converts that opaque period into a timeline, and it is one of the few things an operator can put in front of an underwriter that materially affects both terms and the retroactive date. The mechanics are set out on agentinsured.eu, on retroactive dates and prior acts. How certification evidence enters underwriting more generally is at how certification feeds underwriting.

The finding that surprises people

Operators preparing for assessment often try to reduce their change frequency in advance, on the assumption that a stable system reads better than a busy one. It does not, and the instinct is worth correcting because acting on it makes the result worse.

An agent whose system instruction was edited forty times in a quarter, each edit dated, attributed, tested and reversible, describes an organisation that is watching its system and responding to what it sees. An agent edited twice with no record describes an organisation that is not watching. The first is a better risk on every dimension that matters, including to an insurer, and it scores accordingly.

What an assessment is measuring is not stability. It is whether the organisation's account of its own system is true, and whether it will still be true next month. Change control is simply the mechanism by which that stays the case, which is why it is scored inside Governance and the Autonomy Envelope rather than as a category of its own. It is not a property of the organisation. It is the thing that makes every other stated property checkable.

Questions

Why does a certification assessment care about prompt changes?

Because for an agent built on a general-purpose model, the prompt is most of the system. The code routes requests and renders output. The behaviour that creates the exposure, what the agent will say, what it will refuse, what it will act on without asking, is set by the system instruction, the tool permissions and the model version. An assessment that certifies behaviour it cannot pin to a version has certified a snapshot with no shelf life. The change record is what makes a certification statement mean anything a week after it is issued.

Does frequent prompt changing hurt an assessment score?

No, and this is the most common misunderstanding. Frequency is not the finding, unrecorded change is. An agent whose system instruction was edited forty times in a quarter, each with a dated entry, a stated reason, a named authoriser and a regression run before release, scores well. An agent edited twice with no record scores poorly. High change velocity with good control reads as an organisation actively managing its system. Low velocity with no control reads as an organisation that does not know what its system currently says.

What if the model provider changes the model and we changed nothing?

This is a real and frequently unmanaged exposure, treated as a control question rather than a fault. Two answers score. The first is version pinning where the provider offers it, with a documented process for evaluating and adopting a new version rather than being moved onto it. The second, where pinning is unavailable, is a scheduled regression suite that would detect a behaviour change, with results retained. An organisation that can say neither cannot state what its agent does today, only what it did when someone last looked.

What does a minimum viable change record contain?

Six fields per entry are enough for a small team: the date and time the change reached production, which surface changed, the previous and new value or a stored diff reference, who authorised it, what test was run before release, and how it would be rolled back. Held in version control alongside the application, this takes minutes per change and is the single artefact that most improves both an assessment result and an insurance submission. It is also what makes an incident investigable.

Which dimensions is change control scored in?

Primarily Governance and the Autonomy Envelope, with regression and testing evidence read into Performance and Reliability. It is not scored as a dimension of its own because it is not an independent property of an organisation. It is the mechanism that makes every other claim about the agent verifiable over time. A strong statement about human oversight is only as durable as the process that stops those controls being edited away without anyone recording it.

Sources

  1. European Parliament and Council. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence, Articles 12, 14, 26 and 72. OJ L, 12.7.2024.
  2. Regulation (EU) 2026/1744, the Digital Omnibus on AI. Entered into force 27 July 2026. Annex III obligations from 2 December 2027, Annex I from 2 August 2028. European Commission, AI Omnibus enters into force, checked 17 August 2026.
  3. European Commission, AI Act Service Desk. Timeline for implementation of the EU AI Act, checked 17 August 2026. Source for the 2 August 2026, 2 December 2027 and 2 August 2028 dates.
  4. International Organization for Standardization. ISO/IEC 42001:2023, Information technology, Artificial intelligence, Management system. Geneva, 2023.
  5. National Institute of Standards and Technology. AI Risk Management Framework 1.0 (NIST AI 100-1), released 26 January 2023, functions GOVERN, MAP, MEASURE and MANAGE. nist.gov, checked 17 August 2026.
  6. Agent Certified. Methodology specification, published at agentcertified.eu/methodology. Source for the dimensions referenced above and for the evidence tiers applied in assessment.