← All pieces

Everyone Bought AI. Almost Nobody Bought Results.

Nine in ten organisations now use AI — yet only about 6% report significant earnings impact from it, and while 64% of OECD countries run generative AI, fewer than a third measure whether it worked. Here is what separates the two groups, and how Kaizen Spark Tech builds for the second one.

Written by KaizenSpark Admin
A split illustration: on the left, a stack of disconnected software dashboards and pilot reports; on the right, a single clean operational workflow with a measured outcome dial, connected by an arrow
A split illustration: on the left, a stack of disconnected software dashboards and pilot reports; on the right, a single clean operational workflow with a measured outcome dial, connected by an arrow

Two numbers from the same survey, published in August 2026, tell you where this industry actually stands.

The first: nearly nine in ten respondents now report their organisation using AI regularly in at least one business function, and 44% say they are scaling it across the enterprise. The second: the share attributing any EBIT impact to AI is 37% — essentially unchanged from a year earlier — and the "high performers," those attributing 5% or more of earnings to AI and describing the impact as significant, still make up about 6% of all respondents. Flat. Year over year. (McKinsey, The State of AI in 2026)

Adoption went vertical. Results went sideways.

We are Kaizen Spark Tech. We build AI automation and agents, marketing and lead systems, web and software platforms, and we advise the people running them. And we would rather open with that uncomfortable statistic than with a promise, because the gap it describes is our entire business.


The reality check nobody selling AI wants to print

Gartner predicts that over 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. In the same release, Gartner names the practice of "agent washing" — rebranding chatbots, assistants and RPA as agents — and estimates that of the thousands of vendors claiming agentic AI capability, only around 130 are real. (Gartner, June 2025)

This is not a hypothetical risk. In March 2024 the U.S. Securities and Exchange Commission brought what were widely described as its first AI-washing enforcement actions, settling charges against two investment advisers for $400,000 in combined civil penalties over AI capabilities they did not have. (SEC)

Meanwhile Deloitte's 2026 enterprise survey of 3,235 business and IT leaders found that only 21% say their organisation has a mature governance model for agentic AI — meaning roughly four in five are scaling agents faster than the guardrails around them. (Deloitte Insights)

If you are a buyer, that is the market you are shopping in. If you are a vendor worth hiring, that is the market you have to prove you are not part of.


What actually works — with receipts

The failure statistics are only half the story, and the honest half is that AI does work, reliably, in a specific and narrow set of conditions.

The strongest evidence in the field is peer-reviewed. Brynjolfsson, Li and Raymond studied a staggered rollout of a generative AI assistant across 5,179 customer support agents and measured a 14% average increase in issues resolved per hour — and a 34% improvement for novice and low-skilled workers, with minimal effect on the most experienced staff. Customer sentiment improved. So did employee retention. (NBER Working Paper 31161; published in the Quarterly Journal of Economics, 2025)

Read that finding carefully, because it contains the whole thesis: AI did not replace the workforce. It compressed the distance between a new hire and a veteran.

The public sector has produced the clearest published numbers anywhere — largely because the UK government insists on publishing them:

  • Consult, an AI tool for analysing consultation responses, processed over 50,000 responses for the Independent Water Commission, categorising them into themes in around two hours at a cost of £240, with experts needing just 22 hours to verify. It agreed with one or both human expert groups almost 83% of the time — while two well-practised human groups agreed with each other only 55% of the time. (DSIT, October 2025)
  • Extract, a planning-data tool, turns records that a planning officer might spend up to two hours on into work done in two minutes for some document types, with only minor edits needed in around two thirds of cases. (MHCLG Digital, June 2026)
  • A cross-government Microsoft 365 Copilot trial across more than 20,000 civil servants found an average self-reported saving of 26 minutes per person per day. (GOV.UK, June 2025)
  • In Singapore, GovTech's OneService chatbot handles over 30,000 cases a month, saving roughly 2,000 staff hours and cutting resolution times by up to two working days — without additional headcount. (GovTech Singapore, January 2024)
  • Iceland's Askur assistant addresses 90% of citizen correspondence, materially reducing calls and emails to service centres. (OECD Digital Government Outlook 2026)

Now the part most vendors edit out. In the same UK programme, the Redbox assistant reached 5,330 officials at its peak — and development was discontinued (the code was open-sourced, and commercial tools overtook it). In a trial of AI coding assistants across 1,000+ staff in 50 departments, staff saved the equivalent of 28 working days a year, and only 15% of AI-generated code was used without any edits. (GOV.UK, September 2025)

And the U.S. Census Bureau, surveying working Americans in March 2026, found that among those who had used AI at work in the previous week, 10% said it saved them no time at all and 3% said it cost them time. (U.S. Census Bureau)

Any vendor who cannot tell you which of their deployments landed in that 13% is either not measuring, or not telling.


The single variable that separates winners from pilots

Across all of this evidence, one factor keeps reappearing.

McKinsey found that the AI high performers were distinguished not by model choice, not by budget, but by workflow: nearly three-quarters of high performers had fundamentally redesigned workflows because of AI — up from 55% the year before — versus roughly a quarter of everyone else. (McKinsey)

This is why so many pilots die. An AI agent bolted onto an unchanged process does not remove the process; it adds a layer to it. The organisation now has the old workflow and a tool, plus the cost of both.

Government has documented the same failure from the procurement side, with brutal clarity. The UK National Audit Office examined five large digital change programmes and found that commercial approaches to working with suppliers contributed to delays totalling at least 29 years and more than £3 billion in cost increases — at least 26% above original forecast (the NAO notes this may not reflect final cost). Its diagnosis names the mechanism precisely: frameworks "risk encouraging a contracting approach based on time usage rather than specifying what the supplier is meant to achieve… geared to buying inputs rather than outcomes." (NAO, January 2025)

The U.S. picture is structurally identical. GAO has kept federal IT acquisition on its High-Risk List since 2015, noting that the government invests more than $100 billion a year in IT while investments "too frequently fail or incur cost overruns and schedule slippages while contributing little to mission-related outcomes." Of 69 legacy systems reviewed, the 11 most in need of modernisation ranged from 23 to 60 years old, and seven were operating with known cybersecurity vulnerabilities. (GAO-25-107852, GAO-25-107795)

Buy inputs, get invoices. Buy outcomes, get outcomes.


How we work: four stages, and a stop button at every one

1. Diagnose before we propose. We do not lead with a product. We map where work actually queues — the inbox nobody owns, the re-keying between two systems, the after-hours enquiries that go cold. Then we quantify it. If the honest answer is that automation will not pay for itself, we say so, and you keep your money.

2. Redesign the workflow, then automate it. This is the step that produced the 3× difference between McKinsey's high performers and everyone else, and it is the step most vendors skip because it is unglamorous and cannot be demoed.

3. Ship narrow, measure hard, then scale. One process. A defined baseline captured before go-live. An agreed metric — resolution time, cost per case, conversion rate — reviewed on a fixed cadence. Notably, the U.S. federal AI acquisition memo now requires exactly this discipline of its vendors: contracts must let agencies evaluate performance "e.g., on a quarterly or biannual basis," and evaluation data "should not be accessible to the vendor." (OMB M-25-22) We think that is a good standard for private clients too.

4. Hand over the keys. Documentation, training, and no architecture that holds your data hostage. M-25-22 makes vendor lock-in protection a required contract term for U.S. federal buyers; we apply it as a default everywhere.


For public sector buyers: what we bring to the table

Governments are not short of AI. They are short of evidence. The OECD found that 64% of OECD countries use generative AI for at least one purpose — but that only 28% report any financial or non-financial measurement of the impact of their AI use cases. Adoption without evaluation is how a decade of "transformation" produces a decade of pilots. (OECD Digital Government Outlook 2026)

We build to be procurement-ready, not procurement-surprised:

  • European Union. The AI Act (Regulation (EU) 2024/1689) reached its general date of application on 2 August 2026, with transparency obligations now enforced — its prohibitions have applied since February 2025 and its general-purpose AI obligations since August 2025. Note what many vendor sites still get wrong: the AI Omnibus, Regulation (EU) 2026/1744, deferred Annex III high-risk obligations to 2 December 2027 and Annex I to 2 August 2028. We classify every system against the Act's risk tiers before a line of code is written. (European Commission; Regulation (EU) 2026/1744)
  • United Kingdom. We build to the Algorithmic Transparency Recording Standard, mandated across all government departments on 6 February 2024 — which states plainly that commercial suppliers selling algorithmic solutions to public bodies "should be comfortable with this level of transparency that is expected of the public sector." We are. (GOV.UK) Buyers should also note that Crown Commercial Service became the Government Commercial Agency on 1 April 2026.
  • United States. We design against the NIST AI Risk Management Framework, Section 508 accessibility requirements built into the requirements document rather than retrofitted, and the contract terms required under OMB M-25-22 — including its provision that contracts permanently prohibit the use of non-public agency data and outputs to further train commercially available AI models absent explicit agency consent.
  • India. The DPDP Rules, 2025 were notified on 14 November 2025 with an eighteen-month phased compliance window, activating the Act's penalties of up to ₹250 crore for failure to maintain reasonable security safeguards — the era of treating DPDP as "an Act awaiting rules" is over. We build to MeitY's India AI Governance Guidelines, including their expectation of grievance redressal and published transparency reporting. (PIB / MeitY)

Three questions worth asking anyone

If you run a business, an agency or a department and you are being told AI will transform it, three questions will tell you most of what you need to know about who you are dealing with:

What baseline will you measure against? What workflow are you changing, not just automating? And which of your deployments did not work?

The third one matters most. Anyone who has done this at scale has a deployment that underdelivered, and can tell you precisely why. Anyone who hasn't will change the subject.

We answer all three before a contract exists — which usually starts with mapping a single process end to end, quantifying what it currently costs in hours and money, and saying plainly whether automation is the right answer for it. Sometimes it isn't, and that is a useful finding to reach before the budget is committed rather than after.

[Talk to us about a process worth examining →](#)


Kaizen Spark Tech builds AI automation and agents, marketing and lead-generation systems, web and software platforms, and provides operational consulting for private and public sector organisations. Every statistic in this article is linked to its original published source. If you find one we've got wrong, tell us and we'll correct it in public.

handoverdeliverydocumentation

If you have inherited a system with no documentation and need this produced retrospectively, we do that as a standalone engagement, including for software somebody else built.

Keep reading

Related notes

When automation is the wrong answer

Three situations where we tell clients to fix the process first and quote for less work.

In review

Ask what you would be handed.

Before you sign with anyone, ask them to describe their handover pack. The quality of the answer tells you what kind of relationship you are entering.