How to Run an AI Pilot That Survives Contact With Production
The UK government ran a rigorous, 1,000-person trial of AI coding assistants and published the uncomfortable result alongside the good one. That is what a trustworthy pilot looks like.

You have probably seen the claim that 95% of AI pilots fail to reach production. It traces to a single 2025 working paper — a preliminary report, not peer-reviewed, based on 52 interviews — and the figure has since been repeated across dozens of blogs with invented precision it never had. We are not going to add to that pile. What follows instead is drawn from one of the few AI trials run at real scale with a disclosed method.
The trial worth learning from
The UK government ran an AI coding assistant trial across more than 1,000 staff in 50 departments. The result, reported plainly: engineers saved the equivalent of 28 working days a year, 72% agreed the tools were good value, and 65% completed tasks faster.
Then the same report stated the uncomfortable part without softening it: only 15% of AI-generated code was used without any edits. (GOV.UK)
That is what a trustworthy pilot looks like. It measured a real benefit and reported a real limitation in the same release, at a scale large enough that neither number is noise.
What made it trustworthy
Scale. Over 1,000 participants across 50 departments is large enough that the results are not one enthusiastic team's experience.
A named, checkable baseline. "28 days a year" and "15% used unedited" are specific enough to be wrong if they were fabricated, which is exactly what makes them credible.
It published the bad news voluntarily. Nobody made the UK government report the 15% figure. A pilot report that only contains good news should be read with that absence in mind.
Designing your own pilot this way
Fix the baseline before you start, not after. Measure the current process — time, cost, error rate — before the AI system touches it. A pilot that only measures the "after" state cannot tell you what changed.
Pick a metric that can embarrass you. The UK trial's 15%-unedited figure is not flattering, and its presence is what makes the 28-days figure believable. If every metric in your pilot plan makes the system look good, you have not designed a pilot. You have designed a demo.
Run it long enough to see the tail, not just the average. A single good week proves nothing. The failure modes that matter — the edge case nobody anticipated, the input format nobody tested — tend to appear later, not in week one.
Decide the kill criteria before you start, not when you're emotionally invested in a yes. What result would tell you to stop? Write it down before the pilot begins, because the same organisation that would honestly answer that question in week one will rationalise a bad result by week twelve.
Measure who it helps, not just whether it helps. The peer-reviewed evidence on AI in customer service found the average benefit was 14% — and 34% for novice staff, with almost no effect on veterans. (Brynjolfsson, Li & Raymond, QJE, 2025) A pilot that only reports the average would have missed the most useful finding in that entire study.
What to do with an ambiguous result
Most honest pilots are ambiguous. That is not a reason to hide the result — it's a reason to say so plainly and either scope the next stage narrower or stop. The UK's own AI programme is instructive here too: the Redbox assistant reached 5,330 officials at its peak and was discontinued anyway, its code open-sourced, once better commercial tools existed. Nobody in that programme pretended the story was uncomplicated, and the programme is more credible for it.
If you are designing a pilot and want someone to stress-test the metrics before you commit budget to scaling it, that is a conversation we are glad to have.
Kaizen Spark Tech designs and delivers software, AI, automation and digital infrastructure for businesses and institutions. Every statistic here is linked to its original published source. Where a widely-repeated pilot-failure statistic could not be traced to a credible methodology, we have said so rather than repeated it.
