Generate 200 lines of code in a few seconds, then spend ten minutes checking whether they really belong in the codebase.

If you use AI a lot for development, that probably sounds familiar.

For the past two years, the question of AI productivity has often been summed up pretty simply: AI writes code faster, so developers must work faster too. In its most enthusiastic version, the job disappears along the way. But writing code is only one part of the work.

You still have to understand what to build, work within an existing architecture, review the diff, test, debug, review the change, and decide whether you are confident enough to ship it to production.

And once you measure that whole loop instead of the number of lines produced, things get a little more complicated.

The study where developers were 19% slower

In early 2025, METR ran an experiment with 16 experienced open-source developers across 246 real tasks.

These were not beginners asking ChatGPT to make a TODO app. The participants worked on repositories they had known for several years.

For each task, AI use was randomly allowed or forbidden. Before the experiment, developers expected to save about 24% of their time. In the end, they were 19% slower with AI. That is the figure people remember, even though the context really matters.

The experiment took place in early 2025. Cursor was widely used, mainly with Claude 3.5 and 3.7 Sonnet, and the workflow still looked a lot like this: generate, read, fix, prompt again, repeat.

Domenic Denicola, one of the participants, later explained that he had set up neither project-specific Cursor Rules nor a custom MCP server. For the nine AI-assisted tasks, investing time in that infrastructure simply did not seem worthwhile to him. That makes perfect sense in the context of the experiment.

But it also means the study measured something very specific: expert developers working in their own codebases with the tools available in early 2025.

It did not measure a 2026 Codex or Claude Code workflow with repository instructions, skills, MCP, tests, debugging tools, and agents working on their own for several minutes. METR emphasizes this point.

Why slow down someone who already knows their project?

This is probably the most interesting part of the study. These developers had a huge amount of implicit context.

They knew the architecture, the project conventions, the fragile parts, some historical decisions, and probably plenty of small details that had never been written down. The model arrived without all that knowledge. Asking it to help could therefore create extra work before it saved any.

Fewer than 44% of Cursor generations were accepted. And accepting a generation did not mean taking it as-is: 75% of participants said they reviewed every line, and 56% regularly had to make substantial changes. Producing code faster is not that useful if someone still has to check all of it.

This is also something I have noticed more and more in my own use of agents: the problem is not necessarily that the model cannot code. It first needs to know how this project should be coded, and that context does not appear out of thin air.

Not all tasks are the same

Not every METR participant was slowed down either.

Quentin Anthony, who took part in the study, later explained that some tasks worked pretty well with AI: documentation, tests, and certain refactors. Others worked much less well.

His rule was simple enough: if the model is not moving toward a correct solution after about ten minutes, take over.

That helps avoid a pattern I know a little too well. The model is almost there. One more prompt, then another. Twenty minutes later, I am not really debugging my code anymore. I am debugging my conversation with the agent.

Knowing when to use AI, and especially when to stop using it, is part of the job too.

METR tries again in 2026

METR then launched a larger experiment.

By February 2026, it included 57 developers, 143 repositories, and more than 800 tasks. The median participant had about ten years of experience.

This time, the estimates pointed in a much more positive direction: about an 18% speed gain for people who had taken part in the first experiment and 4% for new participants.

So did AI go from making developers 19% slower to 18% faster in a year?

Not really. METR specifically says the data does not support that conclusion. The confidence intervals are wide, and, more importantly, the way people used AI had started to change.

Some developers no longer wanted to take part if they had to do half their tasks without AI. Others avoided submitting important tasks because they were afraid those tasks might end up in the no-AI group.

And several were now running agents in parallel or doing something else while an agent worked.

At that point, measuring time already gets strange.

If Codex is working on one feature while Claude Code handles another and I review the diff from a third agent, how much time did I actually spend on each task?

So METR eventually revised its protocol.

Faster, yes. More productive?

METR tried another approach by analyzing 5,305 Claude Code transcripts produced by seven technical staff members at the organization in January 2026.

Depending on the method used, some tasks appeared to be completed roughly 1.5 to 13 times faster. That would make a great YouTube thumbnail.

But even METR says we should not translate that into “developers are thirteen times more productive.”

Users naturally choose tasks where they think AI will be useful. And a task that has become very cheap may now get done even though it would simply have been abandoned before. So people may be doing more things. That does not mean the organization is producing thirteen times more value.

The analysis also has limits: seven people, one month of data, and an estimate of the time without AI that partly relies on an LLM judge.

This is a good example of the problem I had in mind when I started this article.

What exactly are we measuring when we talk about productivity? Time to produce the code? Time until there is a mergeable PR? Time until the feature is in production?

You can implement something five times faster and still spend exactly as long waiting on the spec, review, or QA.

The work is shifting

The 2025 DORA report reaches a similar observation by another route.

Nearly 5,000 technology professionals were surveyed. Around 90% said they used AI at work, and more than 80% felt it improved their productivity.

A DORA analysis published in 2026, based on open-ended responses from Google engineers, adds quite a bit of nuance to that impression.

AI speeds up initial generation and makes it easier to get started on a task. Then some of the saved time comes back elsewhere: prompting, checking, and auditing the code. DORA calls this the verification tax.

And above all, the workload may simply shift. A developer can now generate much more code. The reviewer has not suddenly gained four brains and two extra hours in the day. The PR just arrives sooner.

DORA also observes that greater AI adoption is associated with more delivery throughput, but also with more delivery instability. Their idea is that AI works more like an amplifier.

A team with good tests, clear APIs, an understandable architecture, and good tools can produce more quickly.

A poorly documented codebase with few tests can also produce technical debt much more quickly.

And that is where the question of the harness starts to become much more interesting than the question of which model is being used.

There are real gains too

Fortunately, not every story is about developers spending all their time checking Claude’s hallucinations.

A study published in Management Science combines three randomized experiments at Microsoft, Accenture, and a Fortune 100 company, for a total of 4,867 developers.

Access to GitHub Copilot was associated with 26.08% more completed tasks in the combined analysis.

Less experienced developers seemed to see greater gains too. That matters: in real professional settings, AI assistance can increase the amount of work completed. But the study primarily measures the use of a code-completion assistant.

It does not say releases shipped 26% faster or that software quality improved by 26%. Once again, it depends on what we choose to measure.

It is also worth keeping in mind that two authors were affiliated with Microsoft Research, and Microsoft owns GitHub and Copilot. That does not invalidate the results of a randomized experiment, but it is useful to know who funds and produces the research we cite.

And sometimes, nothing happens

In 2024, Uplevel published a study of around 800 developers using GitHub Copilot.

This time, there was no notable improvement across several efficiency metrics. The company even reported a 41% increase in the bug rate for the group with access to Copilot.

But this was not a randomized experiment like the METR study or the one above. Uplevel compared data from before and after Copilot adoption, as well as different groups of developers. So we cannot conclude that Copilot caused 41% more bugs.

The study also measured access to the tool, not exactly how it was used. In short, the result is interesting, but the conclusion is much more limited than the neat “41% more bugs with AI” headline.

When you combine all the studies

A meta-analysis published as a preprint in May 2026 combined 23 studies conducted between 2019 and 2025.

The overall effect on productivity was positive, with Hedges’ g at 0.33. That does not mean “33% more productivity.”

Hedges’ g measures a standardized effect size. Using commonly cited benchmarks, 0.33 corresponds to a modest positive effect.

Above all, the results differed enormously from one study to another. Gains were generally higher in controlled experiments and lower in professional or open-source environments. That is not very surprising.

“Implement this function” is a fairly easy task to isolate. “Find out why this screen sometimes keeps stale state after returning from the background, understand which layer owns that state, fix it without breaking other flows, and prepare a PR the team will accept” is a little harder. Yet both end up in the same category: “software development.”

Green tests, rejected PR

METR also published a study in March 2026 that I find especially interesting.

They took AI-generated PRs that passed SWE-bench Verified and gave them to real maintainers.

Four maintainers from three projects took part.

The automated evaluation was, on average, 24.2 percentage points more optimistic than the maintainers about whether the changes could be merged.

About half of the PRs that passed the benchmark would not have been merged as-is.

That does not mean the agents were incapable of finishing the work. They were not allowed to update their PR after receiving feedback from the maintainer.

But it is a reminder of something basic:

Tests passing does not mean the work is finished.

The code can work and still be wrong for the project: a poor abstraction, an odd API, duplicated code.

So, more productive?

Yes, no, it depends.

The average effect of AI tools seems fairly positive today. Some tasks can even be sped up dramatically. But the closer the studies get to real work in a codebase, the more variable the results become.

What interests me most is what sits around the model.

How much do we know about the project? What instructions have we given it? Can the tests really catch a bad change? Can the agent build, launch the app, read the logs, and understand an error?

Are we asking it to change 80 files at once, or giving it a task small enough that a human can still understand what happened?

In short: are we just using a very good model, or have we built an environment where it can actually work?

That is also why directly comparing Cursor in early 2025 with an agentic workflow in 2026 is becoming difficult. The models have changed, but above all, everything we build around them has changed.

The experiment I would like to see next

If I could choose the next experiment, I would take experienced developers, have them work in their own repositories, and give them real tasks.

But instead of comparing “AI / no AI,” I would compare two ways of using it. On one side, a classic coding assistant. On the other, a current agent with repository-specific instructions, project-specific skills, MCP where useful, build and debugging tools, tests, CI, and a real review loop.

And above all, I would not measure only the time needed to produce a patch.

I would measure the time until a PR is actually accepted, review effort, regressions, and bugs that show up later. Because that may be where the way we measure AI is falling behind the way we use it.

I can generate much more code today than I could two years ago. The hard part is still knowing whether I should have generated that code in the first place.

Sources