Post-market monitoring (AI Act)
What is post-market monitoring under the AI Act?
Post-market monitoring is the duty to keep watching a high-risk AI system after it is in use, instead of checking it once before go-live and assuming it stays the way you left it. Article 3, point 25 of the AI Act defines it as all the activities a provider carries out "to collect and review experience gained from the use of AI systems they place on the market or put into service", so that anything needing correction gets noticed in time.
The duty is split over two companies. Article 72 gives the provider a monitoring system and a written plan. Article 26(5) gives the deployer, the company using the system under its own authority, the duty to watch how it runs and to tell the provider what it sees. Article 26(6) adds the part companies eventually get asked about: keep the logs the system generates automatically, for at least six months. Which role you are in is decided per system, and the provider and deployer entry works through that test.
It does not touch every AI tool you own, because all of this hangs on the high-risk classification. A chatbot on your website is out of scope for Article 72. A system that ranks job applicants is not. The high-risk regime itself moved: Regulation (EU) 2026/1744, the Digital Omnibus on AI, in force since 27 July 2026, pushed stand-alone Annex III systems to 2 December 2027 and AI built into products under Annex I to 2 August 2028.
Why a check before go-live cannot be the whole story
Ordinary product safety works because a conveyor belt built to spec keeps behaving like a conveyor belt. An AI system does not give you that. What it does depends on the data it meets, and the data it meets is the world, which moves. So a system can pass its assessment honestly and be worse six months later while nobody has touched a line of it. The mix of people applying changed, a product range was added, a form field was renamed and half the inputs now arrive empty. The model answers with the same confidence it had on day one, which is what makes the decay hard to see from outside. Model drift and data drift are the entries for the mechanism itself.
Recital 155 says the monitoring system is there so that risks from systems which continue to learn after being placed on the market can be addressed in time. A system that keeps updating itself is the sharp version. A frozen model in a moving world is the common one in a Belgian company.
What the provider has to build
Article 72(1) asks the provider to establish and document a monitoring system proportionate to the technology and to the risks. Paragraph 2 says what it does: "actively and systematically collect, document and analyse relevant data" on how the system performs across its lifetime, data that may come from deployers or from other sources. Actively rules out waiting for complaints, and the point of it is to judge whether the system still meets the requirements it was assessed against in the first place: data quality, accuracy, human oversight, logging. Paragraph 3 turns that into a document, the post-market monitoring plan, which is part of the technical documentation in Annex IV.
The Act told the Commission to adopt a template for that plan by 2 February 2026. The Commission's own AI Act page, last updated 3 August 2026, lists the guidelines and codes of practice it has published, and no such template is among them. As a buyer you can still ask your vendor for the plan, or at least for the part saying which data it expects from you and how. A vendor who cannot answer has probably not written one.
What you have to do as a deployer
Article 26(5) is one sentence: "Deployers shall monitor the operation of the high-risk AI system on the basis of the instructions for use and, where relevant, inform providers in accordance with Article 72." Two things follow from it that are easy to miss.
The first is that the instructions for use are the yardstick, not your own idea of good performance. So read them and pull out what the provider says the system is for, where it is weak, and which measures it expects you to run. If they say nothing on the subject, that gap belongs to the vendor. The second is that you are a data source for someone else's monitoring, because what you see the provider mostly cannot: which cases your people overruled and why, which complaints came in, which outputs nobody ever used.
Then Article 26(6): keep the logs the system generates automatically, to the extent they are under your control, for a period appropriate to the intended purpose and at least six months. Six months is a floor, not a target. If a decision the system helps make can be contested a year later, and hiring and credit decisions can be, six months of logs will not answer the question you get asked. Providers carry the same duty in Article 19, with the same minimum. And "to the extent such logs are under their control" is the phrase to watch with cloud software: the logs sit at the vendor, so who can export them and what happens the day you stop paying are contract questions.
Monitoring and incident reporting are not the same duty
Monitoring runs continuously and mostly produces nothing dramatic. Reporting is what happens when something crosses a line. Article 73 sets the thresholds and the clocks for serious incidents, and the AI incident entry works through what counts, who reports and how fast. The connection worth holding on to: monitoring is what lets you notice, and the logs are what let you say who was affected once you have.
What the monitoring looks at
Article 12 says a high-risk system has to allow for the automatic recording of events over its lifetime, and names what those logs are for: spotting situations where the system presents a risk or has been substantially modified, feeding the provider's monitoring, and letting you monitor operation. That is the shape of what you watch.
Does it still do what the documentation says. Measure against the accuracy the provider claims, not a target you invented.
Has the population shifted. The output can look fine while the input has quietly stopped resembling what the system was built for.
What do people overrule. Overrides are the cheapest signal you will get, because a human already did the analysis. Zero over six months usually means nobody is really looking, which the human oversight entry goes into.
What are people complaining about. Article 85 lets anyone complain to the market surveillance authority, and a complaint you handled yourself is one that never goes there.
Every incident and near miss, including the ones that stayed inside the building.
Most of that overlaps with what a company already runs. With LLM observability or an AgentOps setup, inputs, outputs, latency and cost are already logged. If a data team runs drift checks, half the population question is answered, and support's complaints are written down somewhere already. The work is rarely building something new. It is deciding which of these is the record, keeping it long enough, and putting a name on the person who reads it.
Post-market monitoring versus the conformity assessment
The two get mentioned in the same breath and answer different questions. What separates them is when in the system's life each one happens.
The conformity assessment is a gate. It happens before the system is placed on the market, it is the provider's job, it judges the system at one moment, and it ends in an EU declaration of conformity and CE marking. Passing is a condition for selling. Post-market monitoring is a duty that runs. It starts where the gate ends and lasts the lifetime of the system, and Article 72(2) names its purpose as evaluating continuous compliance with those same requirements. Same yardstick, different moment. A substantial modification sends the system back through a fresh assessment, and the monitoring is how you notice one has happened.
For a deployer the translation is short: the assessment is something you check the vendor did, the monitoring is something you do yourself, for as long as the system is switched on.
Asked for evidence, a year later
A logistics company with ninety people uses a tool that scores applications for warehouse and driver vacancies. Recruitment sits in Annex III, so the system is high-risk and the company is the deployer. In March a candidate is rejected. In February of the following year her lawyer writes, and the question is fair: on what basis, and did a person look at it.
The answer has to come out of records, and the list is longer than it looks. The log entry for that application, with what went in, what score came out and when. The version of the system running that day, because the vendor updated the model in June. The record that the HR manager was the assigned overseer in March and reviewed this file rather than the batch, whether she overruled the score, and if she did not, what she was shown when she agreed with it. The instructions for use in the version that applied then. And the notice the candidate was given telling her a system was involved, which Article 26(11) requires.
Eleven months on, the six-month floor has long passed. A company that kept nothing beyond the minimum has no March logs, sees only the current model version in the vendor portal, and its honest answer is that it does not know. None of that is a technical failure. It is a retention setting nobody chose.
The version that works is not much heavier. The tool exports its logs monthly into a backed-up folder, kept for as long as a claim about the decision could still arrive, which for a hiring decision is years rather than months, and each export carries the version number the vendor published. The reviewer's decision is one line: agreed, overruled or set aside, with a reason in a few words. Setting that up is a morning of work.
Which is why the decision worth making first is not which metric to track. It is where the logs live and how long they are kept. A metric you did not track, you can start tracking tomorrow. A month you did not keep is gone.