Digital Economy Dispatch #301 -- Is Your AI Deployment Running on Faith?

We have moved from seeing pilot trials of Microsoft Copilot to organisation-wide deployment in under two years. The evidence underneath that journey has not travelled nearly as far.

In almost every conversation I have with public and private sector organisations about AI deployment, Microsoft Copilot is the answer before the question has been asked. The reason is rarely capability.

Ask what Copilot can do and you will get a fluent, often enthusiastic answer. People describe the meeting summaries, the email triage, the first drafts that no longer start from a blank page. But ask why this tool rather than Gemini, or ChatGPT Enterprise, or something built around the specific work the organisation actually does, and the conversation moves briskly away from capability and towards procurement.

For most organisations, Microsoft Copilot is often the path of least resistance. It sits inside an agreement they already have in place. It requires no new supplier onboarding, no fresh security review of an unfamiliar vendor, no separate business case for a product the board has never heard of. Press a little harder and the honest version usually emerges, which is that it was the only option that could realistically be delivered inside the current structural and financial constraints.

I have considerable sympathy with that. When the alternative is a twelve-month procurement exercise, avoiding adding a line to an existing enterprise agreement is not laziness. It is the only route to making progress. But it is worth being clear about what has happened. The tool was not chosen because it was the best fit for the work. It was chosen because it was the one that could be most readily obtained. That is a perfectly reasonable way to pick a starting point for AI deployment. And a much weaker way to deliver on your AI strategy.

Consider what may be the largest AI deployment in the UK, or anywhere in the world. In June, NHS England announced that 505,000 staff will receive Microsoft 365 Copilot, with 200,000 onboarded in the first six months. That sits inside a five-year programme that Microsoft says will eventually reach 1.5 million NHS employees.

What it costs has not been disclosed. At Microsoft's own UK list price of £23.10 per user per month, 505,000 licences would come to roughly £140 million a year, and 1.5 million to something over £400 million. Large public sector buyers never pay list price, so the real figure is certainly lower and possibly far lower, but as The Register noted, a deployment of this size is worth well into nine figures annually on public pricing. We have been told what the benefit is, in minutes per day. We have not been told what the bill is.

A commitment on that scale invites a simple question: what evidence do we have that this works, and that it is worth what it costs?

How Did We Get Here?

From my experience, the pattern of AI deployment in most organisations over the last two years has followed a similar pattern. A department runs a pilot. The pilot reports enthusiastic users and a headline figure for time saved. That figure becomes the justification for the next, larger commitment, whose own evaluation then justifies the one after it.

We can see this in action. Barclays trialled Copilot with 15,000 staff and extended it to 100,000. Lloyds Banking Group reached nearly 30,000 licences with a reported 93 per cent active usage. The NHS commitment rests on a pilot across 90 organisations and 30,000 staff that reported, according to Computer Weekly, an average saving of 43 minutes per person per day.

Each step is larger than the last, and each is justified by the satisfaction of the people who took the previous one. However, is this enough? Four serious independent evaluations of Microsoft Copilot deployment exist, and they are remarkably honest about what works and what doesn’t.

The Government Digital Service cross-government trial covered 20,000 licences across twelve organisations including DWP, HMRC, the Home Office and the Welsh Government. It gathered 7,115 survey responses and usage data from 14,500 users and reported an average saving of 26 minutes per day. The Department for Business and Trade went further methodologically, using a diary study rather than a retrospective survey. Australia's Digital Transformation Agency commissioned Nous Group to evaluate a whole-of-government trial of 5,765 licences, and the Australian Treasury published a detailed evaluation of its own 218-person trial.

Read them and a consistent picture emerges. Using Microsoft Copilot for meeting transcription and summaries works everywhere, without exception. Email triage works. Using it to create first drafts of documents is valued. Yet, generating Excel spreadsheets with Microsoft Copilot is poor in every single study, and PowerPoint generation produced no measurable time saving at all in the DBT diary data. The clearest and most consistent benefit, mentioned in three of the four, went to neurodivergent colleagues, disabled staff, part-time workers and people for whom English is a second language.

The evaluations were also candid about their own limits. The GDS report states plainly that its time savings are self-reported estimates, that users were uncertain what the saved time was actually used for, and that such figures should be read alongside wider literature on what these metrics are worth. The Australian Treasury found that only 22 per cent of participants used the tool four or five times a week, that staff judged it inferior to commercially available alternatives because of the security restrictions their own organisation had imposed, and that 65 per cent of managers reported no impact on the quality of work at all.

These limitations need highlighting again. When the people doing the work report time savings and the people supervising it see no change in output, you are looking at something other than a straightforward productivity gain.

Where Do the Numbers Come From?

Every headline figure I have quoted, 26 minutes, 43 minutes, 46 minutes, comes from asking people how much time they think they saved. Not one comes from measuring what they did.

Fortunately, somebody has measured. Microsoft's own research group ran a randomised controlled trial across 66 large firms and 7,137 knowledge workers, allocating licences by lottery and taking every measurement from Microsoft 365 telemetry rather than from surveys.

The results are sobering. Email time fell by 1.4 hours per week, around 12 per cent, and that effect is robust. Time spent authoring documents in Word showed no statistically significant change. Meeting time and attendance in Teams showed no significant change. Out-of-hours work fell by around fifteen minutes a week. After correcting for multiple testing, most of the document and meeting effects disappear entirely.

While these are disappointing numbers, we must be careful not to lean on this too hard. The trial ran on an early version of the product during its pilot rollout, and Copilot in 2026 is much improved from Copilot in 2024. Take-up was modest, with the median participant using it in only 39 per cent of the weeks they had access. The authors are candid that their Word analysis was underpowered, that organisational adaptation to a new technology usually takes longer than six months, and, most importantly, that telemetry cannot see quality. They observed no work content and no performance measures at all, so a better document written in the same two hours is invisible to them. The participating firms also opted in, and the highest-priority users were excluded from the lottery by design.

Take all of that seriously and the finding should be viewed carefully. But it doesn’t disappear. What the study establishes is that when several thousand knowledge workers were handed Copilot at random, the only change in behaviour large enough to detect was in how they dealt with email. That is a genuine benefit, and I wouldn’t dismiss it. It is simply a much smaller claim than the one currently carrying major deployments such as half a million NHS licences.

The wider labour market evidence points the same way. Though it covers AI chatbots in general rather than Copilot. Anders Humlum of Chicago Booth and Emilie Vestergaard surveyed around 25,000 Danish workers from 7,000 workplaces in each of two rounds, linked to national administrative records running to June 2024. Users reported average time savings of 2.8 per cent of working hours. On pay and hours, the authors report what economists call precise zeros: not that they failed to find an effect, but that the data rules out any effect larger than one per cent. "We could not detect a change" and "we can be confident there was no meaningful change" are not the same finding, and only the second tells you anything.

They also found that chatbots created new tasks for 8.4 per cent of workers, including people who never used them. Saved time does not automatically become available time.

Some evidence points the other way. When the research group METR randomly assigned 246 real tasks across sixteen experienced open-source developers, those permitted to use AI tools took 19 per cent longer to finish. The sample is small and the setting particular, and the authors are careful not to generalise. But one detail travels anywhere: the developers predicted AI would speed them up by 24 per cent, and afterwards still believed it had, by 20 per cent. The stopwatch disagreed with all of them. Australia's evaluation found a milder version of the same mechanism in Copilot itself, where time spent checking inaccurate output ate into the time the tool had saved.

Understanding the Boundaries

Put the two halves together and the real failure becomes visible. The evaluations were honest. The announcements were not dishonest either. What happened in between is that the limitations and qualifications from these studies fell away.

The 26 minutes travelled upward through briefing notes, business cases and ministerial statements. The sentence sitting next to it, the one saying the figure is self-reported and should be treated with care, did not. By the time a number reaches a press release it has been stripped of the conditions that made it meaningful, and a carefully hedged research finding has become a procurement justification.

I don’t think anyone in that chain is behaving badly. Self-reported satisfaction is the only thing that can be measured cheaply at the scale of a departmental trial, and it happens to be the measure that vendors, programme owners and ministers all have reason to find persuasive. The incentive structure reliably produces this class of evidence. Nobody has to bend anything.

The challenge we face is that you cannot generate deployment-scale evidence without deploying at scale. The measurement problem is clear: an assistant that saves fifteen minutes here and there across half a million people produces effects too diffuse for any pilot to detect cleanly. Waiting for better evidence is not a neutral act either, because the counterfactual is not a controlled experiment. It is staff using consumer AI tools on personal accounts with no governance at all. Faced with that, deploying a governed tool inside an existing security boundary is a defensible bet.

So, the objection is not that NHS England is wrong. It may well turn out to be right. The objection is that a bet of this size is being placed as though it were a calculation, and the organisations placing it are not building the instruments that would tell them, in two years, whether it paid off.

Delivering AI Value

When deploying AI in any organisation, three questions must be faced.

First, what will we measure, and how? If the answer is a satisfaction survey at month six, you will learn how people feel, not what changed. Agree now what telemetry or workflow measure you will look at, before anyone has a stake in the result.

Second, where did the saved time go? Every evaluation I have read leaves this unanswered, and it is the only question that determines whether minutes saved become value delivered.

Third, are we deploying an assistant or an agent, and does our evidence cover the one we are actually buying?

None of this argues for stopping your AI deployment. It argues for knowing what you are doing. There is no shame in running on faith when the evidence cannot yet exist, but there is real risk in mistaking faith for arithmetic, and being unable to tell the difference.