When AI Eats the Control Group
~12 min read · By Jeroen Janssen · September 2026
Domenic Denicola has maintained jsdom since 2012. It is a JavaScript implementation of a web browser, over a million lines of code, and for most of that time he has been its only active maintainer. In early 2025 he assembled nineteen pieces of his own work, each small enough to finish in under two hours, and handed control of them to a randomiser. For each item, a coin decided whether he was allowed to use AI on his own codebase.
Fifteen other experienced developers did the same. Across 246 tasks, being allowed to use early-2025 AI tools made the work take 19 percent longer.
Sixteen developers is a small study, and the authors were asked about that. They replied that they do compute confidence intervals accounting for the number of developers, but that these were “not reported in the released paper, but forthcoming” (METR, 2025). So the 19 percent went out without them.
That is worth noticing, because this is the good end of the evidence. A randomised trial, forecasts collected in advance, a result the researchers had no reason to want. If the best number in this argument reached the public without its error bars, nothing on a vendor’s slide is going to be better.
The number is not the interesting part. This is: before starting, the developers expected AI to make them 24 percent faster. Afterwards, having done the work and lived through it, they still believed it had made them about 20 percent faster. Economists asked to predict the study said 39 percent faster. Machine-learning researchers said 38 percent (Becker et al., 2025).
Sixteen expert maintainers working in codebases they had known for five years are not every developer, every task or every tool. The result was widely used as though they were.
Then METR ran it again, larger, and something stranger happened.
The follow-up covered 57 developers, more than 800 tasks and 143 repositories. The estimated slowdown shrank, to 18 percent among returning developers and 4 percent among new recruits, and the researchers declined to believe their own numbers. Between 30 and 50 percent of participants told them they were holding back tasks they did not want assigned to the no-AI condition. Developers who were unwilling to work without AI at all had declined to join. And people had stopped being able to say how long anything took, because while an agent worked they worked on something else (Becker et al., 2026).
The experiment was losing the conditions under which its question could be asked. Nobody would sit still long enough to be the control group.
None of that means the effect cannot be measured. It was measured, carefully, and the answer was uncomfortable. It means the window for measuring it cleanly is closing, and closing before it produced anything settled.
The usual fallback when measurement fails is the judgement of experienced people. This study removed that one first. Sixteen expert maintainers worked through a nineteen percent slowdown and came out believing they had been twenty percent faster. Nobody in that picture had a reason to be wrong. They simply could not tell.
So a business case for this technology has neither the instrument nor the intuition to lean on. It has to stand on something else, and the rest of this is about what that something else has to be.
Why I care what the number says
I have written these cases, and defended them.
Twenty-five years inside public institutions and delivery organisations, starting as a developer and ending up responsible for the governance around the delivery machine rather than the machine itself. Today I advise on AI governance and risk, teach the subject, and run a small applied research practice for stress-testing governance positions before reality does it for them.
So I have sat on both sides of the table in rooms where a number in a very large font survived contact with everything that should have killed it. I have also watched governance become the reason nothing happened. That is the interest you should weigh what follows against.
Start with what delivery is
Before you can say what AI costs, you have to say what it delivered.
The least misleading unit I have found is an accepted state change inside a named value stream. A requirement becomes an operational capability. An incident becomes a restored service. Prompts, tokens, hours, tickets and merged pull requests show that something happened. None of them is the change that was promised.
The practical difficulty sits between systems. A tool like IBM Apptio brings technology spend into a portfolio view. A tool like Tempo Financial Manager derives labour cost from time booked against Jira work. One lives near the investment decision, the other near the work. The join between them is the interesting part, and it is usually missing: which people and which machines contributed to which accepted change (IBM Apptio, n.d.; Tempo Software, n.d.).
The fix starts as plumbing. Give contributions a work identifier. Attach licences, inference, integration, security review and rework. Connect the work to a deployment and the deployment to an outcome. Then divide.
All of that is worth doing. None of it makes the resulting percentage true.
Why one number cannot travel
Productivity is discussed as though it were a property of the tool, something the vendor ships and you install. It is a relationship between a tool, a person, a task, a work system, a time period and whatever the alternative was. Change one of those and the sign can flip.
Brynjolfsson, Li and Raymond studied a generative assistant used by 5,172 customer support agents across three million chats. Issues resolved per hour rose 15 percent on average, and by 36 percent in the lowest skill quintile. The most experienced gained little speed, and some of their quality measures fell (Brynjolfsson, Li and Raymond, 2025).
The average was not wrong. It was incomplete in exactly the way a board should care about. The assistant appears to have carried the better agents’ practice to the people with the most to learn. Among experts doing unfamiliar work, or in a system where the constraint is testing and approval rather than typing, there is no reason for that average to walk through the door with the licence.
Minus 19 percent and plus 15 percent are both real. They only look like a contradiction once someone removes the people, the work and the system from the denominator.
Where the money goes missing
Say an assistant frees six developer hours. Those hours reach the accounts through one of three routes: output you sell, spending you avoid, or a delay you remove that someone was paying for. Otherwise you have capacity, and capacity is not cash.
They may not even stay hours. The developer picks up another request. Reviewers inherit more to read. An afternoon’s prototype turns into a month of hardening. Everyone stays busy and a slide announces a saving.
The denominator has to include what no supplier invoices: learning time, error review, controls, exceptions, rework, context switching, and the pilots that went nowhere. Model consumption is visible because someone bills for it. Human attention arrives disguised as ordinary work. The cost that is easiest to count is usually the smallest one.
Then there is the number everybody cites. The claim that 95 percent of enterprise generative AI initiatives return nothing came from a preliminary 2025 MIT Project NANDA report built on 300-odd public initiatives, 52 interviews and 153 surveyed leaders. Deployment figures were directional, definitions of success varied, and several effects were reported by participants rather than independently measured (Challapally et al., 2025).
The report may well have found something real. It did not establish that 95 percent failed in any single comparable sense. The figure travelled as a settled verdict while the evidence underneath it was still contested.
A design that would hold
Name the accepted change and record where you are starting from. Attach the full human and machine cost. Keep quality, delay, failure and distribution visible rather than collapsing them into one figure. Use randomisation while it is still credible, phased rollout or matched cohorts when it is not, and an explicit causal argument when no clean comparison survives at all. Write down in advance what result would kill the case.
Then put a date on it, and make the sponsor stand in front of the board with three verbs. Scale, when the outcome improved against a credible alternative with full cost and quality counted, and when you know where the benefit is going. Repair, when the uncertainty has a bounded cause and one more finite test. Otherwise stop.
I thought that solved it. Better evidence would rescue the business case from the evangelists and the undertakers at the same time.
Two words in my own design would not sit still.
Accepted by whom, full for whom
Accepted by whom. Full for whom.
Who decides what counts as delivery. Who decides which displaced burden goes into the denominator and which one is allowed to leave the building. Who decides whether freed capacity becomes more service, more profit, more learning, shorter hours or fewer people.
Every business case has a denominator, and every denominator has a border. Costs pushed past it do not stop existing. They land on suppliers, on the workers in those suppliers, on communities near the infrastructure, on public institutions, on the environment, and on the people the system gets wrong.
Crawford (2021, pp. 18-19) draws the perimeter wider than any ledger: AI is “an idea, an infrastructure, an industry, a form of exercising power, and a way of seeing.” It is material, resting on capital, labour, extraction and planetary logistics. A firm that counts licences, inference and review time has costed the part that reaches its own accounts.
Those outside costs are not one kind of thing. A supplier’s electricity is a priced input. Water stress absent from that price is an externality. Supplier concentration is a dependency. Discrimination may breach a right before anyone gets near pricing it. Forcing all of that into euros manufactures a comparability that does not exist.
Nor does every effect belong on a work item. The accepted state change is the unit of delivery. Energy use and supplier concentration may only be measurable per provider or per portfolio. The unit of delivery and the boundary of responsibility are different things. Good accounting connects the two levels. Bad accounting invents an allocation.
So the business case is missing an account of who is authorised to draw its border.
Who may draw it
In Who Gets to Decide? I set out six powers: to See, Speak, Decide, Act, Stop and Share (Janssen, 2026). I expected to add them to the equation as another input. That was the wrong picture. They govern who may build the equation, who may contest it, who may authorise it, and who lives with the answer.
See and Speak govern the evidence. People have to be able to tell measurement from inference, find the costs that were moved rather than removed, and correct the picture. An engineer who challenges a throughput number that ignores maintainability has to get an answer, not an acknowledgement. So does someone carrying a burden that sits outside the project ledger. Permission to complain is not the power to change the claim.
Decide and Act connect the claim to authority. The sponsor defends the delegation, the board authorises its limits, and the record shows what actually reached the world in the organisation’s name. Usage is not delivery, and nobody can answer for an action that cannot be reconstructed.
Stop and Share keep success from sealing itself shut. A pilot needs a decision date, a loss limit, a quality threshold, a fallback, and a named person who is protected when they invoke any of them. “We are still learning” cannot become a permanent exemption. And a saving has not arrived anywhere when it enters a benefits register. Someone has to decide where the capacity, the profit, the workload, the environmental burden and the risk landed. Pilot purgatory is full of experiments nobody was ever authorised to end.
None of this recreates a control group or proves causality. What it does is narrower: it stops a causal claim from acquiring authority while staying immune to challenge.
The case against all of this
The serious objection grants the analysis and calls it a luxury.
The argument runs like this. Your competitors are adopting these tools now. The evidence will not settle for years, and by the time it does the question will have changed. Every previous general-purpose technology was adopted long before anyone could measure its return, and the firms that waited for proof lost. A stage-gate process with pre-registered kill criteria, boundary analysis and six named powers is how a large organisation converts a two-week experiment into a two-quarter governance programme. Meanwhile a competitor with worse analysis and better nerve ships.
Most of that is right. There is a real cost to rigour, and it falls disproportionately on the people trying to build something.
Where the argument fails is in what it assumes about reversibility. Moving fast on a reversible bet is cheap and correct. The rigour is for the bets that are not reversible: the ones that restructure a department, remove the people who know how the old process worked, sign a multi-year commitment, or put a system between a company and the people it serves. The burden should rise with scale, dependency, opacity, irreversibility, and the distance between whoever gets the gain and whoever carries the risk. A governance process that treats a spelling suggestion like an automated refusal has stopped being careful and started decorating.
What would show I am wrong
The harder objection is that the six powers are memorable language for competent management rather than anything that predicts a return. Perhaps well-run organisations both exercise them and invest well, and the framework explains nothing at all.
That is testable. Across comparable portfolios, initiatives showing the six conditions should have smaller gaps between forecast and realisation, fewer material costs discovered after scaling, and earlier termination of the losing ones. If they do not, this is ceremony. If organisations keep producing answerable outcomes while one of the six is systematically absent, the claim that they cannot substitute for each other is weaker than I have made it.
I have not run that study. Nobody has.
As these tools become infrastructure, the credible alternative will stop being “the same work without AI.” It will be another model, a different division of labour, a narrower delegation, or not doing the work. The counterfactual is going to keep weakening. The burden of proof should not weaken with it. And a positive internal return does not settle a case whose costs left through the boundary.
Monday
Take one live AI initiative. Ask the sponsor and the delivery lead to put on a single page: the accepted change, the credible alternative, the internal cost, the material external effects, the quality and legitimacy thresholds, the review date, who holds the authority to stop it, and where the released capacity is going. Connect AI consumption to the work items and the work items to an operational outcome. Then let the people doing the work, receiving it, and carrying its consequences correct what the page says.
If the page cannot be filled in, do not scale. That empty page is the first finding.
An ordinary business case authorises a future. An AI business case has to stay alive long enough to withdraw that authorisation, so that the organisation can still change its mind before contracts, targets and expectations harden around a mistake.
Denicola could find out what the tools did to his work because he ran an experiment on himself for months, in a codebase he has maintained for over a decade, with a randomiser deciding what he was allowed to use. Almost no organisation will do that, and the ones that try will find, as METR did, that people will not hold still to be measured.
By February 2026 he was back in the same codebase with an agent, spending three weeks rebuilding its resource loading while on holiday in Japan, and his view of the tools had moved. “Although back in ye olde July 2025, AI agents on average slowed down experienced developers working on large codebases, these days they’re probably a speedup” (Denicola, 2026). The developer the clock caught out now believes he is faster, and he may well be right.
What he does not believe is that this settles anything. He called that piece “The Wrong Work, Done Beautifully,” and his conclusion reaches something no ledger in this essay does: the agents “haven’t solved the need to plan and prioritize and project-manage.” His worry is a team that gains real speed and spends it badly, shipping on the original schedule “with beautifully-extensible internal architecture, all P3 bugs fixed.”
That is the failure a fully instrumented business case cannot see. Every accepted change recorded, every cost attached, the arithmetic clean, and the whole apparatus counting work nobody needed. The unit of delivery assumes someone has already decided the change was worth making, which puts the argument back where it started. Accepted by whom.
The only control left is an institution still able to find out, in time, that it was doing the wrong work beautifully.
References
- Becker, J., Rush, N., Barnes, E. and Rein, D. (2025) ‘Measuring the impact of early-2025 AI on experienced open-source developer productivity’, arXiv, 2507.09089. Available at: arxiv.org/abs/2507.09089 (Accessed: 2 September 2026).
- Becker, J., Rush, N., Cunningham, T., Rein, D. and Mahamud, K. (2026) ‘We are changing our developer productivity experiment design’, METR, 24 February. Available at: metr.org (Accessed: 2 September 2026).
- Brynjolfsson, E., Li, D. and Raymond, L. (2025) ‘Generative AI at work’, The Quarterly Journal of Economics, 140(2), pp. 889-942. Available at: academic.oup.com (Accessed: 2 September 2026).
- Challapally, A., Pease, C., Raskar, R. and Chari, P. (2025) The GenAI divide: State of AI in business 2025. Project NANDA, July. Available at: cloudelligent.com (Accessed: 2 September 2026).
- Crawford, K. (2021) Atlas of AI: Power, Politics, and the Planetary Costs of Artificial Intelligence. New Haven, CT: Yale University Press.
- Denicola, D. (2025) ‘My participation in the METR AI productivity study’. Available at: domenic.me (Accessed: 2 September 2026).
- Denicola, D. (2026) ‘The wrong work, done beautifully’, 5 February. Available at: domenic.me (Accessed: 2 September 2026).
- IBM Apptio (n.d.) ‘What is technology business management?’. Available at: apptio.com (Accessed: 2 September 2026).
- Janssen, J. (2026) Who Gets to Decide? Six Powers for Human Work in the Age of AI. Deventer: Apparens. Available at: books.apple.com (Accessed: 2 September 2026).
- METR (2025) ‘Measuring the impact of early-2025 AI on experienced open-source developer productivity’, 10 July. Available at: metr.org (Accessed: 2 September 2026).
- Tempo Software (n.d.) ‘Financial Manager FAQ’. Available at: help.tempo.io (Accessed: 2 September 2026).
