- Digital Economy Dispatches
- Posts
- Digital Economy Dispatch #296 -- Three Paradoxes of AI Adoption
Digital Economy Dispatch #296 -- Three Paradoxes of AI Adoption
Three paradoxes are derailing AI adoption. AI helps most where you are already expert. Supervising it demands more skill, not less. And the productivity gains arrive last, not first.
I’ve spent a lot of time over the past few weeks in discussions with senior leaders about accelerating AI adoption. The same exchange keeps happening. Someone sets out what they want generative AI for and focuses on the work they cannot do themselves. Write the code nobody in the team ever learned to write. Draft the contract clause they would otherwise pay a solicitor for. Digest the hundred-page report nobody has time to read. Each time, I find myself giving the same unwelcome reply.
That instinct is exactly backwards.
It is also the first of three paradoxes that explain a lot about what is currently going wrong in AI adoption. None of them is a technical problem. All three are management problems. And each one has a direct consequence for how you invest, how you train your people, and how you judge whether any of it is working.
Paradox One: AI helps you most with the things you already know how to do
The most rigorous evidence we have on this comes from a field experiment run by Harvard Business School researchers with Boston Consulting Group. Fabrizio Dell'Acqua and colleagues describe what they call a jagged technological frontier: AI assistance improves performance on some tasks and actively worsens it on others, even within the same workflow and at apparently similar levels of difficulty. The experiment was preregistered, ran with 758 consultants, around 7% of BCG's individual contributor population, and has since been published in one of the top academic journals, Organization Science.
The finding that matters for leaders is not the average uplift for the team. It is what happened at the edges. On tasks that fell outside the frontier, knowledge workers using AI performed worse than those working without it. The paper is blunt about why: users tended to over-rely on the tool, and combined human and machine performance fell precisely where closer supervision was needed.
That leads to an uncomfortable implication. The frontier is jagged, which means you cannot tell from the outside which tasks sit on which side of it. The only reliable way to know whether an AI-assisted answer is good is to already know enough to judge it. Expertise is not what AI replaces. Expertise is what makes AI safe to use.
This inverts the usual business case. Organisations tend to deploy AI where capability is thinnest, because that is where the pain is most obvious. That is precisely where it is most dangerous. So, for example, the team with no legal training is the team least equipped to spot a plausible-sounding clause that will not survive contact with a court.
An important counterpoint is that the frontier moves. Ethan Mollick, one of the paper's own co-authors, writes constantly about the ever-expanding range of what these systems can do, and models that failed a task last year may pass it now. That is true, and worth watching. But it does not help the manager deciding what to deploy this quarter, because the frontier's shape is not published anywhere. It has to be discovered locally, by people who know the work.
Paradox Two: the better the tool, the more skill you need to supervise it
This one is not new. It was set out in 1983 by Lisanne Bainbridge, then a psychologist in the Department of Psychology at University College London, in a five-page paper called Ironies of Automation. She was writing about industrial process control and aircraft flight decks, but the argument transfers to knowledge work almost without alteration.
Her point was this. When you automate most of a task, you leave the human responsible for the part that cannot be automated, which is usually the rare and difficult part. But because the human no longer performs the routine version day to day, the underlying skill decays. So, at the exact moment intervention is needed, the person best placed to intervene is least practised at it. Bainbridge's conclusion was that automation means operators need more training, not less.
Put paradox one and paradox two side by side and you have a slow-acting trap. AI works best in the hands of people who already know the work. Sustained use of AI erodes the knowledge that made those people effective. The organisation that treats AI as a substitute for building expertise will find, in three or four years, that it no longer has anyone able to tell whether the output is any good.
For enterprise-scale organisations and public sector leaders, this is the smart-buyer problem in a new guise. You cannot outsource your way to the capability you need in order to judge what you have outsourced.
Paradox Three: the productivity gains arrive last, not first
In July 1987, Robert Solow reviewed a book for the New York Times and produced a line that has outlived almost everything else he wrote for a general audience: you can see the computer age everywhere but in the productivity statistics.
We now have a decent explanation for why. Erik Brynjolfsson, Daniel Rock and Chad Syverson call it the productivity J-curve. General purpose technologies, they argue, require large complementary investments: redesigned business processes, new products and business models, retrained people, and much more. Those investments are intangible, poorly captured in the national accounts, and expensive. So measured productivity dips before it rises. The gains are real, but they are back-loaded, and they only materialise once the organisation has changed shape around the technology.
If that is right, then flat results after eighteen months of AI pilots are not evidence that AI does not work. They are evidence that you are at the bottom of the curve, where you should expect to be.
There is an obvious danger in that argument, and it is worth considering. The J-curve can be used to excuse any disappointing result indefinitely, and I have heard it used exactly that way. What stops it becoming an alibi is the second half of the claim: the gains follow the complementary investment and only follow it. So, the question to ask is not "where are the savings?" but "what have we actually redesigned?". If the honest answer is that people are using a chatbot alongside processes that have not changed in several years, you have not started the investment that produces the return, and the curve owes you nothing.
Why Your Evidence Probably is Not Evidence
There is a fourth finding that cuts across everything above, and it should make anyone relying on self-reported benefits uneasy.
The research organisation METR ran a randomised trial with experienced open-source developers working on their own codebases. Participants expected AI to speed them up by around 24%, and reported afterwards that it had made them roughly 20% faster. Instead, measured against the clock, they were about 19% slower.
That result needs careful handling. METR has since changed the design of its follow-up experiment and reported some evidence of speedup, but with confidence intervals wide enough that the direction is not settled: a point estimate of 18% faster among returning participants, on an interval running from 38% faster to 9% slower, and 4% faster among newly recruited developers, on an interval from 15% faster to 9% slower. Selection effects were part of the problem. Some developers were reluctant to take part if they might be told to work without AI, and some avoided submitting the tasks they most wanted help with. METR now describes the original finding as historical, and so should we.
What survives the revision is the gap between perception and measurement, and that gap is critical. Almost every AI business case I see rests on people telling you they feel more productive. That is not evidence. It is a feeling, and in this study the feeling pointed in the opposite direction to the stopwatch.
Three Questions You Must Address
Reviewing these paradoxes raises questions any leadership team must address.
Where are we deploying AI into capability gaps rather than into capability strengths, and what would it take to reverse that?
What are we doing deliberately to maintain the expertise we will need to supervise the AI systems we are deploying over the next five years?
What is our evidence base for claimed productivity gains with AI, and would it survive rigorous analysis?
Taken together, the three paradoxes point the same way: AI rewards organisations that already understand their own work and punishes those hoping it will stand in for understanding it. The investment that matters is therefore not the tool, but the expertise needed to supervise it and the process redesign that lets any gain reach the numbers. The most important consequence is that difference between an organisation that is measuring its AI benefits and one that is merely feeling them will be obvious in about three years, and by then it will be too late to fix.