A pilot should leave more than a demo. It should tell the organisation what to scale, what to change and what to stop.
The meeting after an AI demonstration is often strangely vague. The room agrees that the result looked promising. Someone suggests another use case. Nobody records what the test proved, what it failed to prove or which decision now follows. Six weeks later, the pilot is still alive because ending it would feel more uncomfortable than extending it.
That ambiguity is expensive. IBM’s current review of AI returns says only around 25% of AI initiatives deliver the expected ROI and 16% have scaled enterprise-wide. Organisations do not need fewer experiments. They need experiments that end in clearer decisions.
I would treat the decision record, not the demo, as the primary output of an AI pilot. It is a short, durable account of the assumption tested, the baseline, the operating boundaries, the evidence and the verdict. Its purpose is not to celebrate the technology. Its purpose is to make the next allocation of money, attention and risk more intelligent.
Begin with an assumption that can lose
‘AI can improve customer service’ is not a testable assumption. It is an aspiration. A useful assumption names a workflow, a user, a baseline, an expected change and a condition that would disprove the idea.
For example: an assisted-response workflow can reduce median first-response time from the current baseline to an agreed target while keeping the material-error and escalation rates below defined limits. The numbers will differ by organisation. The discipline is to state them before the result is visible.
This matters because a pilot with no losing condition will nearly always be called promising. The team may discover something interesting, but it will not know whether the original claim survived.
Keep the baseline boring
The baseline rarely makes an impressive slide. It is the current time, cost, error rate, rework, escalation, adoption or customer outcome before the new system arrives. Without it, a polished demonstration can create enthusiasm without creating a comparison.
A baseline also exposes where the problem actually lives. If delay comes from an approval queue, poor case routing or missing data, a better model may not move the outcome. That is useful knowledge before the organisation builds a larger system around the wrong constraint.
Separate model evidence from workflow evidence
An AI system can perform well in evaluation and still fail in use. I find it helpful to separate three kinds of evidence.
Model evidence asks whether outputs meet the required quality, reliability and safety thresholds.
Workflow evidence asks whether the full sequence of human and system actions improves the target outcome.
Deployment evidence asks whether the organisation can operate the system with acceptable cost, controls, support and adoption.
Blending these categories hides trade-offs. A high-quality output may require a review process that removes the time saving. A fast workflow may produce errors that the organisation cannot accept. A technically successful test may depend on data or infrastructure that cannot be sustained.
Name the boundary and the stop authority
The NIST AI Risk Management Framework is useful here because it treats governance, context, measurement and risk management as connected activities. It calls for documented roles, deployment-context measurement and a decision about whether a system should proceed. It also recognises that a system may need to be disengaged or deactivated when it no longer behaves consistently with its intended use.
A pilot record should therefore name the data boundary, affected users, human review, unacceptable failure, escalation path and person authorised to stop the test. This is not bureaucracy added after innovation. It is what allows a real test to run without pretending that every risk can be discovered in a sandbox.
End with a verdict, not an adjective
‘Promising’ is an adjective. It is not a decision. Every pilot should end with one of three verbs: scale, change or stop.
Scale means the evidence supports the assumption within the stated boundary and the organisation is ready for a controlled next stage. Change means the test revealed a correctable problem and a new assumption will be tested. Stop means the expected value, risk or operating reality does not justify further investment now.
The stop verdict deserves more respect. Teams often extend weak pilots because effort has already been spent, an executive sponsored the idea or the demonstration attracted attention. A documented no can save months of engineering and change-management work. It also leaves evidence that prevents another team from quietly restarting the same experiment under a new name.
Use a five-line decision record
The record does not need to become a long report. A useful version can fit on one page.
- Assumption: What specific workflow claim could be proved wrong?
- Baseline: What happens today, using the outcome and risk measures that matter?
- Boundary: Where, for whom and under which controls was the test valid?
- Evidence: What changed in model, workflow and deployment performance, including exceptions?
- Verdict: Scale, change or stop, decided by whom, with the reason and next review date.
Make the record travel
A decision record matters only if the next team can find and use it. McKinsey’s 2025 global AI survey found that AI high performers were nearly three times as likely as others to redesign workflows fundamentally. That kind of change crosses product, technology, operations, risk, finance and frontline teams. A pilot’s evidence cannot stay with the technical project group.
Store the record where investment and operating decisions are made. Link it to the workflow owner, source data, risk review and follow-on work. Use the same core fields across pilots so patterns become visible. Over time, the organisation learns which types of workflow, data and sponsorship are likely to scale and which warning signs appear early.
This is where the record becomes more than project administration. It becomes organisational memory.
Record what did not change
Teams naturally write down improvements. The more revealing evidence is often what did not move. Response time may fall while escalation stays flat. Output quality may improve while user adoption does not. Recording the unchanged measure prevents a partial success from being presented as a complete one.
The same applies to exceptions. A pilot record should include the cases that needed manual rescue, the users who opted out and the conditions under which the workflow broke. Those details are not footnotes. They define the boundary of the result and tell the next team where a larger rollout is likely to strain.
The real test of a pilot
The purpose of a pilot is not to protect optimism. It is to reduce uncertainty at a reasonable cost. Sometimes the answer will be yes. Sometimes the workflow or control model must change. Sometimes the right decision is to walk away.
Before approving the next AI experiment, I would ask one question: if this pilot ends next month, what decision record will remain? If the answer is only a demo, the test is not yet designed to teach the organisation what it needs to know.



