Firms that adopted AI coding agents saw lines of code per worker rise 30%, but the number of finished tasks did not rise by a statistically significant amount, according to a working paper by Harvard researchers Fiona Chen and James Stratton, which Ars Technica reported on 9 October. The paper finds that code review became the bottleneck: after agent adoption, the time to review a pull request rose 49%.
The paper, “Artificial Intelligence in the Firm: Bottlenecks in Software Production”, uses data from Jellyfish, an engineering analytics platform: about 300 million work events from 725,938 workers at 718 firms, between January 2021 and March 2026. It is labelled a job market paper, current version dated 4 August 2026.
What the Paper Measured
Jellyfish combines data from code version control such as GitHub, issue trackers such as Jira, Google Calendar, HR records and usage of GitHub Copilot, Cursor and Claude Code. The authors compare firms that adopted AI tools earlier with those that adopted later or not at all, a staggered difference-in-differences design that relies on adoption timing.
They study two kinds of tool. AI coding assistants, GitHub Copilot or Cursor, reached about 45% of firms in the sample by April 2024. For AI coding agents, the share of firms that had adopted one rose from near zero in October 2024 to over 95% by January 2026.
Adoption is measured at the firm, not the individual. The authors write that their lower estimates likely reflect this, because the firm-level figure averages over workers who took the tools up unevenly.

Business Pill · VERIFICATION COST
A one-minute explainer of verification cost: why something cheap to make is not cheap to check. It teaches the general idea only and says nothing about any person or organisation in this story.
The key insight: As we read it, the paper separates two things that are often reported as one: how much code a team writes and how much finished software it ships. In the authors’ data, agents moved the first by 30% and the second by no statistically significant amount, and the difference shows up as longer and heavier code review.
More Code
After a firm adopted AI agents, lines of code added rose by 1,494 per worker-month, 30% of the baseline mean of 5,017. Commits rose by 4.57 per worker-month, 20% of a baseline of 22.58, and pull requests by 1.22, 23% of a baseline of 5.34.
AI assistants had smaller effects: 12% for lines of code, 9% for commits and 5% for pull requests, with only the commits effect statistically significant. The authors write that their estimates are consistent with, though on the lower end of, recent studies of AI coding productivity.
Not More Finished Work
To measure output, the paper counts Jira issues resolved, which happens once a body of coding work is deployed to production, and Jira epics, which mark finished projects or features. For agents, issues resolved rose by 0.12 per worker-month on a baseline of 3.67, with a standard error of 0.17, which the authors class as not statistically significant. Epics also showed no significant effect.
The authors write that their estimates “rule out an increase in output of larger than 12% of the baseline mean” for agents, well below the 30% rise in code. They also tested whether issues got larger after adoption, using a model that predicts how long each issue would take, and found no evidence of a change in average issue size.
On jobs, they find no significant employment change: for agents, the estimates rule out a decline greater than 2.9% in overall employment and greater than 13.8% in engineering employment.
Where the Extra Code Went
The paper’s explanation is review. After agent adoption, the average time from a pull request being submitted to being merged rose by 3.45 days, 49% on a baseline of 7.03 days. The share of pull requests with changes requested rose by 0.12 on a baseline of 0.13, which the authors describe as nearly doubling the rate, and review comments per pull request rose 35%.
Firms moved people toward review: the share of workers who comment on or approve a pull request rose by 0.04 on a baseline of 0.29, a 14% increase, while the share of workers writing code did not change significantly. The size of each pull request, measured in lines added, also did not change significantly, so the extra review is not just larger submissions.
One junior engineer quoted in the paper describes it: “People could make a huge number of commits quickly, but code review was still the bottleneck.”

AI Reviewing the AI
The authors checked whether AI review tools remove the bottleneck. By March 2026, nearly 80% of firms in the sample had used AI tools in code review. But only 23.3% of all review comments were generated by AI, and only 10.8% of pull requests received at least one AI-generated comment.
A senior engineer quoted in the paper says: “Even with AI tools, you absolutely still need humans involved in the review process.”
What the Authors Tell Managers
The paper’s conclusion says that measuring the returns to AI requires looking at the whole production process: “Many managers rely on intermediate productivity metrics, such as lines of code or commits; however, these may not accurately capture the true gains from AI adoption.”
It adds that realising the benefits requires complementary organisational investments, because speeding up one stage can move the bottleneck to another. In the authors’ model, AI changes both how much draft code is written and how often it contains bugs, which together set how much review is needed.
The Structural Read
The authors’ explanation is that writing code is one step in a chain. Upstream, features are planned and scoped; downstream, code is reviewed and tested before it is deployed. When one step speeds up, the next one sets the pace.
As we read it, the review findings point the same way: the size of each pull request did not change significantly, yet reviews took longer, drew more comments and more often ended in a request for changes. The authors’ model treats this as a change in the bug rate of draft code, not just its volume.
The paper also tracks AI review tools. Nearly 80% of firms in the sample had used them by March 2026, but AI generated 23.3% of review comments, and humans remained central to review in the authors’ account.
A junior software engineer interviewed for the paper (Chen and Stratton, working paper, 2026)
“People could make a huge number of commits quickly, but code review was still the bottleneck.”
Three Implications
ENGINEERING MANAGERS The authors write that lines of code and commits “may not accurately capture the true gains from AI adoption”, and that measuring the returns means evaluating the whole production process.
REVIEW CAPACITY In the paper’s data, firms responded by drawing more workers into review: the share of workers commenting on or approving pull requests rose 14% after agent adoption.
JOBS The paper finds no significant employment change for agents and calls this a reasonably short-run estimate; it rules out declines greater than 2.9% in overall employment.
The Business Engineer Lens
This story maps onto the Business Engineer framework The Agentic Harness War.
The analysis puts it this way: “The constraint is access to systems, management expectations, workforce skill, and review processes — the complements, not the engine.”
As we read it, Chen and Stratton’s paper puts numbers on one of those complements: in their data, code review is where the extra output from coding agents waits.
What Is Not Established
This is a working paper, and the document does not indicate journal peer review. The data come from Jellyfish clients that opted in to research use. The paper says Jellyfish reviewed the manuscript only to check that it did not disclose proprietary information, and “it was not reviewed for its scientific content”.
The authors call their employment result a reasonably short-run estimate and note that employment may respond slowly. They also say a null effect inside adopting firms could mask job losses at other firms if adopters win business from them.
The Bottom Line
In Chen and Stratton’s data, AI coding agents raised code written per worker by 30%, but finished, deployed work showed no significant rise, with a rise above 12% ruled out, and review time rose 49%. Their conclusion: “In particular, code review becomes the bottleneck.”
95,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
A note on sourcing. On 11 October 2026 we read the full 78-page PDF of Fiona Chen and James Stratton’s working paper, “Artificial Intelligence in the Firm: Bottlenecks in Software Production” (current version 4 August 2026), from which every figure and quotation here is taken. Ars Technica’s 9 October report brought it to our attention. The estimates are the authors’, from data supplied by Jellyfish. Nothing here is a forecast or investment advice.
Sources: Fiona Chen and James Stratton, ‘Artificial Intelligence in the Firm: Bottlenecks in Software Production’ (working paper, current version 4 Aug 2026) · Ars Technica, ‘AI coding agents generate more code, but not more software’ (9 Oct 2026)









