Pilot purgatory and proof of value

What is pilot purgatory?

Pilot purgatory is the state where a pilot keeps running without ever going into production and without ever being stopped. Nobody is against it, nobody owns the decision, and every month somebody renews the licence. The people who were supposed to use it have quietly gone back to the old way, and the project still shows up as green on a slide.

The term comes from manufacturing. In 2018 McKinsey, working with the World Economic Forum on the slow uptake of Industry 4.0, used it for companies that had plenty of pilot activity and no bottom-line result to show for it. Their survey at the time found that 84 percent of the companies asked had been in pilot mode for more than a year, and 28 percent for more than two. In the last few years the same word has attached itself to generative AI, because the pattern repeated itself almost exactly.

A proof of value, usually shortened to PoV, is the alternative to a proof of concept. A proof of concept, or PoC, answers one question: can this work technically? A proof of value answers a different one: does it produce a measurable business outcome under real conditions, and is that outcome big enough to pay for running it? The difference is not the technology. It is that a proof of value has a number, a threshold, an end date and a named person who decides, all agreed before the first day.

Think of it as the difference between test-driving a van and running your delivery round with it for a month while you track fuel, loading time and missed drops. The test drive tells you it starts. The month tells you whether to buy.

Why pilots stall

Pilots rarely die of a technical failure. They die of a missing decision, and the reasons are boringly consistent from one company to the next.

Nobody owns the go or no-go. The pilot was started by an enthusiast, sponsored by a manager and built by a vendor. None of the three is the person who has to sign for the production budget, so the question never gets asked.

No baseline was measured before the pilot. If nobody timed the old process, nobody can say what the new one saved. The pilot ends with a feeling instead of a difference.

Success was defined as "it works". A PoC that extracts the right fields from a sample of invoices has done its job. A business does not run on samples. Without a threshold like "under one minute per invoice", every result can be read as encouraging and none as decisive.

The pilot ran on clean data and the process around it was never touched. The vendor got the tidy PDFs. Production gets the scanned fax from the supplier in Turnhout and the mail with the amount only in the subject line. And the approval step, the exception queue and the handover to accounting still assume a person did the work by hand.

Procurement, IT security and legal only start after the pilot. A pilot on a personal account with a sample export runs within a week. Getting the same tool through a supplier review, a data processing agreement and single sign-on takes months, and by then the sponsor has changed jobs.

The vendor has every reason to keep it alive. A running pilot is a reference and a pipeline entry. A closed pilot with a no is a lost deal. Expect a proposal for a phase two, an extra use case or a wider scope every time the decision comes close. That is their business model, not malice, and it is why the exit rule below has to be yours.

How big the problem is

The most quoted figure comes from The GenAI Divide, the report MIT's NANDA project published in July 2025. Its authors reviewed more than 300 public AI initiatives, interviewed 52 organisations and surveyed 153 senior leaders. Their headline: around 95 percent of enterprise generative AI pilots deliver no measurable impact on the profit and loss account, and about 5 percent capture most of the value. The funnel for task-specific tools is the part that describes purgatory best: 60 percent of organisations evaluated such a system, 20 percent got it to a pilot, 5 percent got it into production. The report also notes that mid-sized companies that succeeded went from pilot to full implementation in about 90 days, where large enterprises took nine months or longer. Being small is an advantage here, if you use it.

Gartner made the second figure that gets quoted. In a press release of 25 June 2025 it predicted that more than 40 percent of agentic AI projects will be cancelled by the end of 2027, and named three causes: escalating costs, unclear business value and inadequate risk controls. Two of those three are exactly what a proof of value settles before the project starts.

Treat both figures as directional. The MIT authors say so themselves, and the definitions of pilot and success differ between surveys. The direction, though, is not in doubt.

Proof of concept versus proof of value

The two get used as synonyms, and that is where a lot of pilots go wrong. The cleanest way to keep them apart is the question each one answers.

Proof of concept: can it work?

A PoC tests feasibility. Can the model read our invoice layouts? Can the agent reach our ERP through the API? Does the accuracy on a sample get above 90 percent? It runs on a sample, often on the vendor's side, and success is measured in accuracy, latency or a demo that does not crash. A PoC is cheap and short, and it should be: it only has to rule out that the idea is technically impossible.

Proof of value: is it worth it?

A PoV tests whether the thing changes a business number. It runs on real data, in your environment, with the people who will use it afterwards, for a fixed period, against a baseline measured beforehand. Its output is a decision, and the criteria for that decision were written down before it started. A PoV can fail on a technology that works perfectly, because the process around it did not change, the exceptions ate the saving, or the cost of running it in production is higher than the gain.

Run a PoC when you honestly do not know whether the technology can do the job. Run a PoV when you already believe it can and the real question is whether to spend money and attention on it. Most AI pilots in 2026 should be the second kind: the models can read an invoice, the open question is what that is worth to you.

How to design a proof of value

A proof of value is a short document before it is a project. Write these eight things down, get the decision-maker to sign them, and only then start building.

  1. A baseline. Measure the current process for two to four weeks before anything changes: minutes per item, error rate, cost per item, backlog. If you cannot measure it now, you will not be able to prove a difference later.

  2. One metric. Pick the number that matters for this process and commit to it. Minutes per invoice, days sales outstanding, first-time-right rate. Secondary metrics can be tracked, but only one decides.

  3. A threshold. The value the metric has to reach for a yes. Write it as a number, not as a direction. Faster is not a threshold; under one minute is.

  4. A fixed end date. Six to twelve weeks is enough for most back-office processes. On that date the decision gets made with whatever data exists. Extending the deadline is itself a form of purgatory.

  5. A named decision-maker. One person, with the budget authority for production, who has agreed in advance to say yes, no or new hypothesis on the end date.

  6. Real data and the real process. Live invoices, live mails, the actual exception queue, the actual approval flow. If a step of the process has to change for the pilot to work, change it during the pilot, not after.

  7. The people who will actually use it. The two clerks who handle invoices today, not the project team. Their adoption is part of the metric: a tool nobody opens saves nothing.

  8. The production run cost, estimated before the pilot. Licences, model usage, the hours somebody will spend on exceptions and maintenance, the security review, the integration work. A PoV compares the measured gain with this estimate. Without it a yes is a guess.

None of this is specific to AI. It is the same discipline as any business case, applied before the money is spent instead of reconstructed after.

A worked example: invoice matching

A distribution company with 1,200 supplier invoices a month wants an AI tool to match invoices to purchase orders and goods receipts. Two people in accounting do this today.

Baseline. Three weeks of timing shows an average of 4 minutes per invoice, including the ones that need a phone call to the supplier. That is 80 hours a month.

Metric and threshold. Average handling time per invoice, measured across all invoices, has to come under 1 minute, and at least 95 percent of invoices have to go straight through without a person touching them. Both conditions, because a fast average with a 70 percent straight-through rate means the tool is only good at the easy ones.

End date and decision-maker. Eight weeks from go-live. The finance manager decides, and has already asked for the production budget line to be prepared so a yes can be executed the same week.

Real conditions. All incoming invoices, including scans and the supplier that mails PDFs with no PO number. The approval flow in the ERP is adapted in week one so that matched invoices skip the manual queue.

Run cost estimate. 350 euro a month in licences and model usage, plus an estimated 6 hours a month of exception handling and maintenance. Roughly 600 euro a month all in.

Result after eight weeks. 96 percent of invoices go straight through in about 20 seconds of automated processing. The remaining 4 percent still take a person about 6 minutes each. Weighted across all invoices that is just over half a minute per invoice, or about 11 hours a month against a baseline of 80. Both conditions are met. The saving of around 69 hours a month is worth a good deal more than 600 euro, so the answer is yes, and the tool goes live in week nine with the same two clerks handling the exception queue.

Had the straight-through rate landed at 85 percent, the average time would still have dropped, but the pilot would have failed its own threshold. That is a documented no, with a clear reason and a number, not a pilot that keeps running because the demo looked good.

The exit rule

Every pilot ends in one of three states, and the decision-maker names one of them on the end date.

  • Production. The threshold was met, the run cost holds, the process is adapted. The pilot becomes the way of working, with an owner, monitoring and a budget line.

  • A documented no. The threshold was not met, or it was met at a cost that does not pay. Write down the metric, the result and the reason, cancel the licence, and revoke the access. A clean no is a good outcome: it cost eight weeks instead of two years.

  • A new pilot with a new hypothesis. The result revealed something the team did not know, for example that the tool works for domestic invoices but not for the foreign ones that cause most of the delay. That is a new proof of value with its own baseline, metric, threshold and end date. It is not an extension of the old one.

There is never a fourth state. "Let's keep it running a bit longer and see" is the doorway to purgatory, and the whole point of writing the rule down beforehand is that nobody has to be the person who says stop on the day.

Two things to watch out for. The first is scope creep dressed up as progress: a second use case proposed before the first one has been decided. Decide the first, then start the second as its own PoV. The second is treating a no as a failure of the team. Return on automation across a portfolio comes from stopping the weak candidates early, and a company that produces clean, documented nos is the one that can afford to try the next thing.

Last Updated: September 3, 2026 Back to Dictionary
Keywords
pilot purgatory proof of value proof of concept PoC PoV return on automation automation debt center of excellence straight-through processing kpi process owner agentic ai