- Retirement is a change to a production environment, not the absence of one. Switching an agent off alters what downstream systems receive, which fallback paths get used and who now does the work by hand.
- Four ways an agent ends and only one is a decision: deliberate retirement, replacement, forced migration when a provider deprecates the model, and abandonment. Most real retirements are the third and fourth.
- Five artefacts should outlive the system: a decision record, a final state snapshot, the logs under a reasoned retention period, a credential revocation record, and an integration inventory.
- The integration inventory is the one that cannot be reconstructed later. Everything else can be rebuilt with effort once the people who built it have gone. That one cannot.
- The certificate lapses with the system and does not transfer to the replacement. The assessment record becomes more useful after retirement, because it is the only contemporaneous account of a system that no longer exists to examine.
- A zombie agent is more dangerous than a live one. A live agent has somebody reading its output, so a provider-side change gets noticed. A zombie has nobody, and it still holds permissions.
The gap is easy to explain and worth naming. Assurance frameworks were written by people thinking about the risk that a system does something wrong while operating. That is the obvious risk and it is real. The risk that a system does something wrong while nominally not operating, or that an organisation cannot account for what a system did once it has gone, is less obvious and arrives later, which means it arrives after the framework was published.
This article sets out what an assessment looks for at end of life. It is a companion to prompt change control, which covers the surfaces that alter an agent's behaviour while it runs, and to maintaining certification after assessment, which covers the period in between. Decommissioning is the third of those and the least written about.
The four endings
Ask an organisation how it retires AI systems and you will be described a process for the first of these. Ask how its last three agents actually stopped running and you will hear about the other three.
Deliberate retirement. Somebody decides the agent has served its purpose. There is a date, a plan, a communication and an owner. This is the case every policy document describes and it is the least common in practice.
Replacement. A new agent takes over the workload. The intention is that the old one is switched off in the same movement. What usually happens is that it is left running alongside, either as a fallback nobody has tested or because a single integration still points at it and disconnecting that integration was a separate ticket. Two systems now perform overlapping work, one of them unowned, and the boundary between them is undocumented.
Forced migration. The underlying model is deprecated by its provider, on the provider's timetable, with a notice period the provider chose. This is now the single most common reason an agent stops being the agent it was, and the important property of it is that it is not a decision your organisation made. It is a deadline your organisation received. The migration that follows is a behavioural change of the largest possible kind, performed under time pressure, frequently without the testing that the original deployment received. A framework built around planned retirement has nothing to say about it.
Abandonment. Usage falls to nothing. Nobody decided anything. The agent still exists, still holds its credentials, still has its tool permissions, and still answers if something calls it. It appears in no review, because reviews are triggered by activity and there is none.
Why the ending is a controlled change
The instinct is that risk decreases when a system is switched off. For the agent itself that is true. For the environment it sat in, a retirement is a change like any other, and it is scored as one.
Something downstream was receiving the agent's output and now receives nothing, or receives a fallback, or receives whatever an integration does when its dependency stops responding. A human process that had been reduced to reviewing the agent's work is now performing that work from scratch, usually with fewer people than it had before the agent arrived, because the headcount followed the automation. A queue that was drained continuously is now drained in batches. None of this appears in a change record if the change record was only ever opened for deployments.
There is also a quieter transfer. Whatever the agent was doing that nobody had noticed it was doing stops happening, and the fact that nobody had noticed means nobody will notice it stopping either. Retirement surfaces undocumented dependencies the way a power cut surfaces which sockets mattered, and the way to avoid discovering them the hard way is an integration inventory produced before the switch rather than after it.
Five artefacts that should outlive the system
An assessment asks for these in this order, and the order is deliberate: it moves from the easiest to reconstruct to the hardest.
1. The decision record
Who decided to retire the system, on what date, on what basis, and what the plan was. Three lines. Where the ending was a forced migration, the record should say so explicitly and name the provider notice that triggered it, because that distinguishes a decision from a response and the two carry different weight in every subsequent conversation. Where the ending was abandonment, there is no decision record and its absence is the finding.
2. The final state snapshot
What the system was on its last day of operation: model and version, configuration, system prompts, retrieval corpus manifest, tool permissions and the scope of each, and the human oversight arrangement in force. Organisations that maintain change control produce this in minutes, because it is the last entry in a series they already keep. Organisations that do not maintain change control cannot produce it at all a month later, and that is the more common outcome.
The snapshot matters because every question anyone asks later takes the form of what was the system doing at the time, and the last recorded state is the anchor for answering it.
3. The logs, under a retention decision
Operational logs are the only account of what the system actually did, as opposed to what it was configured to do. They are also the artefact most likely to be deleted by a default nobody chose, because log retention is usually a platform setting rather than a governance decision.
For high-risk systems the AI Act is specific, and it is worth quoting rather than paraphrasing. Article 12 requires high-risk AI systems to technically allow for the automatic recording of events over the lifetime of the system. Article 19, read at source on 28 August 2026, provides that "Providers of high-risk AI systems shall keep the logs referred to in Article 12(1), automatically generated by their high-risk AI systems, to the extent such logs are under their control", and that "Without prejudice to applicable Union or national law, the logs shall be kept for a period appropriate to the intended purpose of the high-risk AI system, of at least six months, unless provided otherwise in the applicable Union or national law, in particular in Union law on the protection of personal data." A second paragraph places the equivalent duty on providers that are financial institutions within the documentation they already keep under Union financial services law.
Three things follow from that wording and all three are commonly misread. The obligation as drafted sits on the provider, so a deployer's retention duty comes from elsewhere: from its own role obligations, from sector law, and from evidential need. The six months is a floor, expressly subordinate to a period appropriate to the intended purpose, and nobody should read it as a target. And the provision defers explicitly to data protection law, which means the tension between keeping and deleting is written into the Act rather than resolved by it. Which of these obligations attaches to you turns on whether you are a provider or a deployer, a distinction set out at agentliability.eu, on the Article 3 definitions, with the logging duty itself at Article 12 logging and record keeping.
What an assessment looks for is therefore not a particular number of years. It is that a period was chosen deliberately, by a named owner, for a stated reason, balancing three legitimate and opposed pressures: the regulatory record keeping duties that attach to the system in your role, the data protection principle that personal data should not be kept beyond its purpose, and the evidential reality that a claim about something the agent did can arrive long after it stopped running. A reasoned decision that lands anywhere sensible is defensible. A platform default is not a decision at all.
4. The revocation record
Every credential, key, token, service account and tool permission the agent held, with the date each was withdrawn and by whom. This is the most frequently missing artefact in the entire end of life set, and it is the one with the most direct security consequence, because a decommissioned agent that retains valid credentials is an unmonitored account with production access and no user.
The revocation record is also the shortest document in this list. It is a table. Its absence is never a resourcing problem; it is a problem of nobody having been asked for it.
5. The integration inventory
What called the agent, what the agent called, and what was done about each connection at retirement. Inbound and outbound both, because both leave residue: an inbound caller now failing silently, an outbound permission now unused but still granted.
This is the artefact that cannot be reconstructed. The other four can be rebuilt with effort from systems that still exist. An accurate map of what was connected to a system that has been gone for eighteen months lives only in the memory of the people who built the connections, and those people leave. If an organisation produces only one of these five artefacts, this is the one worth insisting on.
What happens to the certificate
Certification attaches to a system in a defined configuration, not to an organisation and not to a use case. When the system is retired the certificate lapses with it, and it does not transfer to the replacement, for the same reason a certificate does not survive an uncontrolled change to a live system. The related question of what happens when the vendor changes rather than the system is treated at whether certification transfers when you switch vendors.
What is worth stating, because organisations get it backwards, is that the assessment record becomes more valuable after retirement rather than less. While the system runs, anyone can examine the system. Once it is gone, the assessment is the only contemporaneous, third party account of what it was, how it was governed and what controls were in place, covering exactly the period that any subsequent claim or enquiry will concern. Filing it as spent is the wrong instinct. It should be retained on the same schedule as the logs and for the same reason.
This is also where the insurance and certification records converge, since a claim arriving years after a system was retired is decided on records nobody was maintaining by then. The claims-made mechanics that determine whether such a claim is covered at all are set out at agentinsured.eu, on retroactive dates and prior acts, and the reason defence cost is driven by the quality of these records is at agentinsured.eu, on defence costs and the limit.
The zombie problem
The fourth ending deserves its own treatment because it inverts the intuition about where risk sits.
A live agent in daily use is watched. Its output goes to somebody who reads it, and when a provider-side model update changes how it behaves, a human notices within days because the answers start reading differently. That surveillance is informal, undocumented and highly effective.
An abandoned agent has none of it. Nothing about the deprecation of its model, the rotation of a dependency, or a change in the corpus it reads from will produce a complaint, because there is no reader. It continues to hold credentials it can act with. If anything calls it, an internal service, a scheduled job, an old integration, it will answer, and it will answer with behaviour nobody has evaluated since the day it stopped being used.
Four questions find them, and they can be asked of every entry in an AI inventory in an afternoon:
- Does this agent have a named owner who would answer an email about it this week?
- Has anything about it changed in the last twelve months, including a model version?
- Are its credentials and tool permissions still valid?
- Is anything still calling it, and does anybody know what?
Two unsatisfactory answers is a finding. Four is a system that should have been retired and instead is simply unattended, which is the worst of both states: none of the value of an operating system and all of the exposure. The inventory discipline that produces answerable versions of these questions is the same one an assessment examines at intake, described at preparing for an assessment.
The retirement runbook
Five steps, in order. None of them is expensive and the whole sequence is a day of work for a system that took months to build.
- Open a change record. Retirement is a change. It gets the same record as a deployment, with the same owner and the same approval, and where it was forced by a provider deprecation the record says so and names the notice.
- Take the final state snapshot before anything is switched off. This is the step most often done afterwards, at which point half of it is gone. Model version, configuration, prompts, corpus manifest, tool permissions, oversight arrangement.
- Produce the integration inventory and act on it. Inbound callers notified, outbound permissions revoked, fallback behaviour for each connection decided and tested rather than discovered.
- Revoke and record. Every credential, key, token and service account, with dates and an actor. Then verify the revocation rather than assuming it, because a revocation that failed silently is indistinguishable from one that worked.
- Set the retention decision and hand it to an owner. Logs, records, the assessment file, and a date on which the decision is reviewed rather than a date on which it silently executes.
An organisation that does this consistently gains something beyond the evidence itself. It can answer, at any moment, how many AI systems it operates, which is a question most organisations cannot answer within an order of magnitude, and the reason they cannot is not that they lost track of what they deployed. It is that they never tracked what they stopped.
Questions
Why does decommissioning an AI agent need governance at all?
Because switching a system off changes the behaviour of everything connected to it, and because the obligations attached to the period it operated do not end when it does. A retirement alters what downstream systems receive, which fallback paths get used, who now performs the work by hand, and which credentials are still valid. Each is a behavioural change to a production environment. Meanwhile the records of what the agent did while it ran are exactly the records a regulator, a claimant or an insurer will ask for, years after the system stopped existing.
What are the four ways an AI agent actually ends?
Deliberate retirement, on a plan with an owner. Replacement, where a new agent takes the workload and the old one is often left running alongside for longer than intended. Forced migration, where the underlying model is deprecated by its provider on the provider's timetable, which is the most common cause and the least planned. And abandonment, where usage falls to nothing, nobody turns it off, and it keeps holding credentials and answering requests. Only the first is a decision. The other three are events that happen to you.
What evidence should survive a decommissioned AI agent?
Five artefacts. A decision record naming who decided, when, on what basis and with what plan. A final state snapshot of model version, configuration, prompts, corpus manifest and tool permissions as at the last day of operation. The operational logs under a deliberate retention decision. A revocation record covering every credential, key, token and permission, with dates. And an integration inventory showing what called the agent, what it called, and what was done about each. The first four are reconstructible with effort. The fifth is not, once the people who built the integrations have gone.
How long should we keep the logs of a retired AI agent?
There is no single number and an assessment does not look for one. It looks for a documented retention decision with a stated period, a named owner and a reason. Article 12 requires high-risk systems to technically allow for the automatic recording of events over the lifetime of the system. Article 19 requires providers to keep those automatically generated logs, to the extent they are under their control, for a period appropriate to the intended purpose and of at least six months, unless applicable Union or national law provides otherwise, in particular data protection law. Three things follow: the duty as drafted sits on the provider, the six months is a floor rather than a target, and the provision defers expressly to data protection law, so the tension between keeping and deleting is written into the Act rather than resolved by it. Against that sits the evidential reality that a claim can arrive long after the system stopped running, when the logs are the only defence.
Does a certification survive when the agent is decommissioned?
The certificate lapses with the system and does not transfer to the replacement, because certification attaches to a specific system in a specific configuration rather than to the organisation or the use case. What survives, and matters more than people expect, is the assessment record. It is dated evidence of what the system was and how it was governed during the period it operated, which is exactly the period any future claim or enquiry will concern. It becomes more useful after retirement, not less, because it is the only contemporaneous account of a system that no longer exists to be examined.
What is a zombie AI agent and why does it matter?
An agent that nobody uses, nobody owns and nobody has switched off. It still holds credentials and tool permissions, still responds when called, and is absent from every review because reviews are triggered by activity. It is more dangerous than an agent in daily use for one reason: a live agent has somebody reading its output, so a provider-side model update gets noticed within days. A zombie has nobody watching, so the same change lands silently on a system that still holds permissions to act. The test is four questions: named owner, any change in twelve months, credentials still valid, anything still calling it.