Digital Economy Dispatch #302 -- What the AI Evidence Can and Cannot Tell You

Academic evidence on AI adoption is rigorous but narrow. Practice evidence is broad but harder to test. The most useful lessons lie where they meet.

In theory, theory and practice are the same. In practice, they’re different.

Most organisations adopting AI at scale are making decisions faster than the evidence can support them. That is not a criticism. Leaders cannot wait for a settled literature before they act, and researchers cannot produce one on demand. However, it does mean that a great deal of money is being committed on the strength of a handful of studies, usually reduced to a single percentage on a slide.

I have been thinking about this while preparing for a Digital Leaders AI Week conversation with Dr Will Venters of LSE, who is also an Amazon Scholar working with the AWS Executives in Residence team. We both sit a little awkwardly across two worlds. We read and write the papers, and we work with the people who have to make the decisions in boardrooms and on factory floors. From that position, one question matters more than any other: how much weight can the evidence bear?

My short answer is that the academic evidence is rigorous but narrow, and the evidence from practice is broad but harder to test. The most useful lessons lie where the two meet.

Rigorous, But Narrow

The best academic work on AI and productivity is very good indeed. Erik Brynjolfsson, Danielle Li, and Lindsey Raymond followed 5,172 customer-support agents and found a 15% rise in issues resolved per hour, with the gains going disproportionately to the least experienced. A Harvard and BCG experiment with 758 consultants found significant improvements on tasks within AI's capability, and a 19 percentage point fall in correct answers on a task just outside it. A randomised trial by METR found that experienced developers took 19% longer with the AI tools of early 2025, while believing they had been about 20% faster.

These are careful studies, and what they show is more interesting than the headlines they generate. The gains are real but uneven. They depend on who is doing the work and on whether the task sits within what the tool does well. People also turn out to be poor judges of their own speed. And the results date quickly: METR's own later update suggests that newer tools probably do speed developers up, though the effect has become harder to measure.

What these studies cannot tell us is what happens to an organisation. They examine one type of task over weeks or months. When Anders Humlum and Emilie Vestergaard linked surveys of some 25,000 Danish workers to administrative records, they found no measurable change in earnings or hours two years on. The tasks may have got faster. Pay and working time did not move.

Broad, But Harder to Test

Evidence of a different kind arrived in late September. AWS published Reimagine: Turning AI into Value, a study Will helped to write, based on nine months of interviews with 154 leaders, including 128 executives deploying AI across 27 countries. It is not a controlled experiment, and it comes from a technology vendor, so it should be read with both points in mind. Even so, it looks directly at the thing the academic studies leave out.

Two findings stood out for me. The first is that the bottleneck moves. AWS describes rebuilding one of its own services with five engineers in seventy days, when the work had been scoped at thirty engineers for a year. Yet as Rahul Pathak, who runs Data and AI Go-to-Market at AWS, put it: "we didn't get code out into production any quicker, even though we wrote it faster." Security review, deployment, and approval had not changed.

The second is that saved time does not turn into value of its own accord. Nearly half of the interviews that described efficiency gains contained no evidence of business impact tied to them. Set that beside the Danish result, and I think they are describing the same thing. Time saved that nobody decides how to use shows up nowhere.

Where the Two Kinds of Evidence Meet

The point where these two bodies of evidence touch is the one I find hardest. The academic studies tell us that AI helps novices most. The Reimagine interviews raise a question that, by the report's own account, nobody could answer. If AI takes over the routine work, and routine work is how people build judgment, how does anyone junior build judgment now? Christina Mack, Chief Science Officer at IQVIA, described seeing AI hand an executive and an intern identical output. The difference lay in what each of them took away from it.

I suspect these are one finding seen from two ends. AI lifts the performance of the least experienced today by supplying what they have not yet learned. Whether they then go on to learn it is a question that a study lasting a few months can’t settle. It is also the question that will decide what our organisations look like in ten years.

Three Things to Do With Imperfect Evidence

None of this is an argument for waiting. It is an argument for using the evidence for what it is good at, and there are three practical steps I would suggest.

First, treat any productivity figure as a hypothesis about a specific task, not a forecast for your organisation. Ask four questions of every study you are shown: what was the task, who were the workers, how long did it run, and who paid for it?

Second, decide in advance where saved time will go and who is accountable for turning it into a result. If nobody is, expect the capacity to be absorbed and the business case to remain unproven.

Third, generate your own evidence. The questions that matter most to digital leaders concern governing agents, redesigning work and moving from pilot to scale, and these are the areas where the research is thinnest. Set a baseline before you deploy and agree what success looks like before you start.

The distance between research and practice is a practical problem, not merely an academic one. Researchers need access to real, messy deployments, including the ones that fail. Organisations need to share more than their success stories. Join us as Will and I explore both sides in our AI Week conversation on 21st October at 2pm.

In the meantime, here is a question for all digital leaders: What is the evidence behind your current AI business case, and would it survive those four questions?