AI coding agents can produce far more code without increasing finished software. A study using 300 million work events across more than 700 companies found agent adoption associated with 30% more lines of code, 20% more commits and 23% more pull requests. Yet completed issues and epics did not rise significantly, while review time increased 49%.
Throughput moved to the review queue
The share of pull requests receiving change requests nearly doubled and comments per pull request rose 35%. Firms responded by increasing the share of workers doing reviews by 14%. Those results describe a production constraint: generating a draft became cheaper, but validating correctness, maintainability and fit with the larger system did not.
METR’s separate research points in the same direction. Its study of SWE-bench-passing pull requests found many would not be merged into a real codebase, showing why benchmark completion and production acceptance are different outcomes. METR also reports that measured productivity effects vary with task, developer familiarity and evaluation design.
The economic metric is accepted change, not tokens or lines
More code can be valuable when it expands experimentation, documentation or test coverage. It can also increase security exposure and future maintenance if reviewers cannot keep up. The relevant denominator is therefore accepted, reliable functionality per engineering dollar—not prompts, generated lines or agent seats.
Companies can improve the equation by restricting agents to well-specified tasks, strengthening automated tests, tracking rework and reserving experienced reviewers for architectural decisions. AI-assisted review may help, but the source found humans still performed most review work through March 2026.
What investors should watch
For software vendors, the study weakens simple claims that code generation automatically produces labor savings. For adopters, it identifies a measurable implementation plan: compare lead time, escaped defects, rework and completed features before and after deployment.
BTI's bottom line
If those do not improve, code volume is activity rather than productivity.
