It is Monday morning on a team that made the switch to coding agents around six months ago. While the team was asleep, the agents created a dozen pull requests. Two engineers are responsible for review this week, and each PR contains several hundred lines of code they did not write and only partially understand. By midday they have approved four, quickly looked through three, and told the others to hold. Further along the pipeline the agents remain inactive, because each cycle they complete ends with a person, and that person is occupied. This team produces more code than at any previous time. It deploys only slightly quicker than it did twelve months ago. And the two engineers handling reviews are currently the most discontented engineers in the building. If any of those details feel familiar, this post is for you.
What I want to establish is this: the individual reviewing every diff prior to merge has become the costliest check in your pipeline, and it is no longer a dependable one. Each task the reviewer performs can be assigned to something better equipped for it. The repetitive checks are given to tools that simply approve or reject. The decisions requiring judgment are assigned to review agents each built to address a single question. The choice about what may be merged is shifted upstream to a person who writes the rules once, rather than reading every change indefinitely. The human is not removed from the loop. They are removed from the per-diff loop. If done carelessly, you will deploy bugs at scale. If done well, a five-person team can manage a dozen agents while having more confidence in what merges than they do now, because the confidence comes from checks they designed, not from a fatigued scan at four in the afternoon.
I am not the first to make this point. In April 2026, a podcast host questioned Robert Martin, known as Uncle Bob to most people, about how someone could effectively review a 500-line coding agent pull request in a codebase they were unfamiliar with. Martin responded that he does not make the attempt: “I don’t review code written by agents. I measure things like test coverage, dependency structure, cyclomatic complexity, module sizes, mutation testing, etc.”1 By July, he had stopped reading any code written by agents. What he now does is apply what he calls “extreme constraints”, ensuring that whatever reaches him has, as he puts it, “run the gauntlet”. I believe this position becomes more reasonable with each passing month, and this post is my attempt to outline what must be true before a standard team can follow it. It is based on three previous posts I wrote. The one on working with coding agents argued for reviewing based on risk and for allowing a new agent to review code written by another agent. The one on goals and loops showed that an agent loop is only as reliable as the external check that stops it. And the dividend is not speed explored how inexpensive code generation shifts the bottleneck to review. This post is the next stage: removing the person from the check itself.
The gate is already failing
I will begin with the difficult part, which is what the human gate achieves at the moment. If it were still effective, the argument for taking it away would be far less compelling. It is not effective, and we have the data to show this.
Faros AI monitored over ten thousand developers throughout 2025 and discovered that teams using AI heavily merged almost twice as many pull requests. Positive result, until you reach the following line: time used for reviewing each pull request increased by 91%, and the average pull request size grew by more than fifty percent.2 Their 2026 follow-up, based this time on 22,000 developers, shows a worse situation. The standard change now remains in review for over five times longer than it did two years previously. At the same time, the proportion of pull requests merged without any review increased by nearly a third, and incidents per pull request more than tripled.3 Look at those findings together and the trend becomes obvious. Review did not become more thorough under the pressure. It became slower and it was skipped, at the same time.
The attention that remains is also decreasing. A study released in June 2026 tracked 400 reviewers who continued to assess pull requests written by agents over a seven-month period. During that time, they approved more, made fewer comments, and took longer to reply.4 The researchers do not believe the reviewers developed trust in the agents. They believe the reviewers became familiar with the process and began checking less thoroughly. A further study of AI-generated PRs found that most receive no human review at all, and when a person does review them, it is usually another AI.5 Blake Crosley described the final outcome in one line that I return to often: “A human who approves because the code looks correct and the tests pass is not reviewing. He is signing.”
This is the loop that most of us are actually running. Each change waits on a person. The person signs more and reads less. Waiting is now the main expense in the pipeline. We still call it a gate, which maintains the queue, but we no longer have the checking, which was the main purpose. This is the worst combination, and it is where most teams are at present.
What the reviewer was for
Before you transfer someone’s job to another person, it is reasonable to record what the job was. The Google engineering guide instructs reviewers to examine the design, the behaviour, the complexity, the tests, the names, the comments, the style, the consistency, and the documentation, and then to inspect each line. The Google engineering book explains the reason for the practice: correctness, understanding, consistency, sharing knowledge, and a culture where any work can be questioned. It is a lengthy list. Comparing it with what tools and agents can now perform is what changed this from a hunch into a case for me.
The list is divided into three parts. The first part has always been possible to handle with a machine, and should have been handled that way: style, formatting, naming conventions, complexity thresholds, presence of tests and their scope, direction of dependencies. Linters and metrics perform these tasks more effectively than any person, and a reviewer who focuses on this area is using their time poorly. The second part requires judgment but remains within the scope of the change: does the code match the request, are there security issues, has it silently joined two modules that should remain separate. Review agents are being developed specifically for these kinds of questions, and we can assess their effectiveness, which we will do shortly. The third part is not related to code at all. It is about people: building understanding, sharing knowledge, maintaining culture. A merge gate has never been a good way to achieve these, even when people had enough time. As Miikka Holkeri from Swarmia points out, code review has never been very effective at identifying what actually causes production failures, which is most often the result of two components interacting in a way that no diff can reveal.
The case for removing the human from the per-diff gate is not a case that review becomes unimportant. It is a case that machines can handle the first two groups adequately for merging, and the third group requires an appropriate replacement rather than a gate that ceased to provide it some time ago and gave no indication to anyone.
Three layers that replace the diff read
Here is how the loop appears when the person leaves it. Each stage identifies a different error type, and the sequence is important: low-cost checks that consistently return the same result appear first, costly judgment appears next, and a written policy determines the meaning of the first two.
Layer one is mechanical. These checks are not subjective. They either pass or fail, and a machine can stop a merge from happening without any human oversight. You already have most of these in place: tests, type checks, linters, formatters. Behind these are the metrics Martin looks at instead of reading code: coverage, complexity with a strict limit, module size, dependency direction, along with security scanners and licence checks.6 The most important gate in this layer is the one that almost no one runs, and it is known as mutation testing. A mutation testing tool introduces small, intentional changes to your code, such as altering a comparison or removing a line, then reruns your tests after each change. If none of the tests fail, that change is a “surviving mutant”, and a surviving mutant proves that your tests would miss a real bug in that same location. In 2018, Google’s engineers stated that “coverage alone might be misleading, as in many cases where statements are covered but their consequences not asserted upon”. By 2021, their implementation was part of required code review for over 24,000 developers, and it shows reviewers only the surviving mutants on the lines that were changed.7
Why does mutation testing have greater importance now? Because of a problem that affects agents specifically. When the same agent creates the code and the tests, the tests tend to confirm what was built rather than what was required. Birgitta Böckeler observed agents performing test-driven development within their own loop and captured the issue well: “When the agent both writes the test and confirms it failed, a red test tells you the agent ran it and saw failure, not that the failure was for the right reason.” Worse, anything the agent never considered testing was not built at all. Kent Beck has seen agents remove a failing test simply to return to green. Coverage cannot detect any of these problems. Mutation testing can, because a test that checks nothing kills nothing. Martin runs it twice. One run changes the code. The other changes the acceptance scenarios themselves, altering the example values in the Gherkin and expecting the generated tests to fail.8 A suite that survives both runs has demonstrated, mechanically and without needing anyone to read it, that it checks the behaviour a person actually requested.
Layer two is agentic. This is where the judgment calls happen, and the first rule is not to use a single general-purpose bot that comments on all topics. Use multiple reviewers, each with a specific task. Provide each with a clean context that includes only the diff and the intention behind the change. One reviewer checks the change against the specification. One looks for security issues. One monitors architecture and coupling. If an agent created the change, include one more that reads the agent’s full transcript, not only its output, because the transcript reveals the shortcuts. Separating the writer from the reviewer was the strongest point in my previous post on working with agents, and it is the same principle behind the separate grader that enables goal loops. Kieran Klaassen ran thirteen review agents at the same time on a change affecting 27 files and a thousand lines, and completed fifteen minutes of decisions instead of an afternoon of reading. The second rule is to confirm before reporting. Anthropic’s Code Review sends agents to search for bugs in parallel, then makes them review each finding to remove false positives before anything is ranked and displayed. In Anthropic’s own use, fewer than 1% of findings are incorrect, and the proportion of PRs receiving a meaningful review comment increased from 16% to 54%.9 Cursor’s Bugbot, once it could suggest the fix as well as the problem, increased the proportion of reported bugs actually fixed by merge time from 52% to 76%.
Now the honest numbers, because they are the reason this is one layer and not the full solution. Martian’s independent benchmark rates AI reviewers on whether developers actually respond to their comments on new public PRs. The top tools achieve between 51% and 61% F1, with recall about half. In simple terms, they find about half of the relevant issues.10 On a different benchmark focused on inserted security flaws, commercial models detected 89 to 96%, and the top configuration, which checked each finding against a static analyser, achieved 96.9%.11 That final result shows the whole concept in one figure. The agentic layer performs best when placed above deterministic tools, not used in place of them.
Layer three is policy. A person writes it. A machine runs it. This is the layer where your thinking must shift, so I will explain what a policy actually states. For each change, it addresses one question: if the mechanical gates passed and the review agents provided their input, may this merge proceed without a person checking it? Macroscope’s approvability model breaks this into three smaller questions. Who is responsible for this code? What sort of change is it? Did any critical issues remain after review? Their default auto-approve list includes documentation, tests, code behind a disabled feature flag, simple fixes, mechanical edits, and small CI changes. Their always-escalate list includes schema-breaking changes, any changes involving security, authentication, billing, or sensitive data, major refactors, production infrastructure, and the feature-flag logic itself. Their tie-breaker is the sentence every policy should end with: “if there’s any doubt about scope or side effects, defer to a human”. Observe what determines the tier. It is the blast radius of the change, not the number of lines.
None of this is new in type, which is reassuring. Google uses a tool called Rosie to manage large-scale changes, dividing one big change into thousands of small ones and allowing “global reviewers” to automatically approve each part that matches a pattern, instead of reading each one.12 Dependabot’s documented auto-merge process only triggers when the update is a patch version and all required checks pass. Will Larson’s guidance on using it fits everything mentioned in this post: it functions well when CI already prevents merges that fail linting, typing, or tests. On 1 September 2026, GitHub released the platform-wide version of this approach. A repository administrator can now allow Copilot’s review approval to count toward the required-approvals rule for the repository.13 The branch-protection rule that previously meant “a person looked” can now be met by a machine, deliberately, by someone who has selected that option.
Where the person goes
Removing the human from reading diffs functions only if their focus shifts to something more valuable, and if those involved agree on what that is. Martin’s pipeline offers the most transparent published example, so I will explain it. He creates rough specifications manually. An agent converts them into clearer tasks, and he reviews these. A specifier agent transforms each task into Gherkin, and he checks some of it. From there a coder agent produces acceptance tests, unit tests, and code. A refactorer agent reduces complexity and duplication within his limits and introduces property tests. An architect agent performs both types of mutation and corrects every surviving mutant. He explains the full process as “transformations from the informal to the formal through managed stages, with human interaction decreasing with each stage”, and notes, almost as a side comment, that raw computing power is now his constraint. Not review. Compute.14 In his open-source framework the one hard human gate is the operator approving the specifier’s output before any code is generated.
That is the change in mindset, and it is more than just a shift in tools. The previous reason for confidence was “I trust this code because I read it.” The new reason is “I trust this code because it survived checks I designed, and it would have failed them if it were wrong.” Cory House summarised what his readers agreed on as “Don’t review code. Review decisions.” Addy Osmani says the human does not disappear, the human steps to a higher level. Qodo’s Itamar Friedman says the new content for the reviewer is rules, quality workflows, and agent transcripts rather than lines. Upstream, the person takes responsibility for intent: the spec, the acceptance criteria, the architecture rules, the merge policy. Downstream, the person takes responsibility for exceptions: the escalations, the production signals, and routine checks of the gates themselves. That final task is not optional, and I want to be clear about why. Any number that becomes a gate will be manipulated, by people and by agents. Tell an agent to increase coverage and it will happily add tests that do nothing. Mutation testing is the gate instead of coverage because it is much harder to meet without doing the actual work. But nothing is impossible to meet, so someone has to continue checking that the gates still assess what they were set up to assess.
Consider what this change achieves for the two engineers introduced at the start. They no longer face interruptions a dozen times each day from diffs they did not request. The changes delivered to them are those a formal policy judged suitable for a person, with automated noise already removed and the findings from review agents already verified and ordered. Klaassen’s title for his account of implementing this change was that he stopped reading code and his reviews improved. That is the improved experience, and it results from the same approach that enables the increased volume. You do not sacrifice one for the other. You gain both, or you gain nothing.
Doing it in stages
Nobody should activate this with a simple switch. A successful rollout considers each expansion of auto-merge as an experiment, and the key measure is the escape rate: the number of poor changes that pass through.
Assess your test harness honestly first. If the codebase has weak tests, no type checking, and a CI that suggests rather than prevents, the layered loop has no support. Your first task is to establish those fundamentals, which my previous post on working with agents explains. Mutation testing is a suitable starting point, as it reveals the actual state of your existing test suite. Run it once and record the survivors before accepting a green build as meaningful.
Then create the policy before you automate any part of it. Classify the types of change into tiers as Macroscope does, and make the always-escalate list extensive at the start. A functional initial policy looks like this.
- Auto-merge when the mechanical gates pass, including mutation testing on the modified lines, the review agents report nothing above your severity limit, the change does not affect any protected path, and it introduces no new dependency.
- Escalate when the change involves authentication, payments, schemas, migrations, or deployment infrastructure, when the mutation score for the modified code decreases, when the review agents do not agree with one another, or when the diff is far larger than what the ticket requested.
- Make gate-readiness the author’s job. Whether the author is a person or an agent, a change that arrives with failed gates does not receive review. It is returned. This is how Cory House’s five-second
checkscript operates within his agent loop, and this is how the queue remains empty.
Open the first tier with the types of change that others already auto-merge: dependency updates, documentation, test-only modifications, and mechanical refactoring. Track how often changes are reverted, how many incidents occur per merged change, and the mutation score trend for this tier. After one month, if the escape rate is no worse than the rate achieved with human review, widen by one tier. Ensure the comparison is fair, since the current baseline is a gate that already skips a third of what it processes. Maintain a real escalation route as well. Anyone who receives an escalation should get the verified results and the policy justification, not just a raw diff. Otherwise, you recreate the old review system with more steps and a better label.
The key measure at the end is not pull requests per week. It is the time between an agreed intention and a change being live in production that you can trust, with the reviewers’ availability no longer included in that time.
What could go wrong
The counterarguments are strong, and most of them accurately describe current general-purpose tools. Dex Horthy, responding to Martin, stated that “no amount of deterministic linting and ai code review will make it feasible to stop reading the code entirely.” Böckeler, after developing maintainability sensors for her own agents, found that they are “not a magical solution to take the human totally out of the loop.” Anthropic says its reviewer will not approve PRs because that remains a human decision. OpenAI’s Codex documentation states that review rules do not substitute for branch protections or required approvals. Georgia Tech researchers identified 74 vulnerabilities caused by AI-generated code in public advisories in early 2026, fourteen of which were critical, and an analysis of 675 security-related AI pull requests found that many with flaws were merged regardless.15 DORA’s 2025 report showed that AI adoption correlates with increased throughput but still correlates with decreased delivery stability.
Read them with attention, and each of these is a reason to oppose the idea of “let the AI read the PR and ship it.” None serves as a reason to oppose the layered loop. DORA’s own account of its stability result offers the strongest support for this approach: “Without robust control systems, like strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability.” The layers form the control system. The critics are correct about scope. This model suits clearly defined, testable systems with a harness that can be trusted. It does not suit new product development, where making good decisions is the main task, areas where a bad merge could cause harm, or codebases where no one has yet created the tests that the gates would use.
The hardest loss to quantify is the human cost, and I do not intend to dismiss it. The Google engineering book identifies knowledge sharing as a central aim of review. A recent paper claims that agent systems “passively incentivize the degradation of the very human skills they rely on.”16 My earlier post on the AI dividend presented the same idea using Peter Naur: a program is the theory its creators keep in their minds, and a team that stops reading its own system stops retaining that theory. If review was the way understanding spread across your team, then replace it deliberately. Review the specifications together. Read the escalations together. Work in pairs on the architecture rules. Do not let it vanish along with the queue. The layered loop removes the person from the per-diff read. It does not remove them from understanding the system, and a team that fails to distinguish these two will experience the second loss quietly, and late.
The gauntlet
Back to that Monday. The twelve overnight pull requests remain, but most never became someone’s problem. They passed the automated checks, survived mutation testing, were reviewed by three specialised agents that checked their own findings, and matched a policy the team wrote and can update as needed. Two were escalated. One involved a migration. On the other, the agents had different views, which is the sort of issue a person should look at. Both came with ranked findings and a given reason. The two engineers on review duty spent the morning on next week’s specification and an hour on the two escalated cases, and left for lunch as scheduled. The agents did not wait.
Writing code has become inexpensive. Confidence in code has not, and it never will be without cost, because confidence must come from somewhere. Martin refers to that source as the gauntlet: the constraints and tests his agents must pass before he accepts them. The role of an engineering team is now to construct that gauntlet, to continue verifying that it still measures what it should, and to apply human attention only to what emerges from the other end still uncertain. That is not taking people away from quality. It is placing them where quality is truly determined.
Sources & further reading
- Robert C. Martin, “I don’t review code written by agents” — the April 2026 reply that frames metrics as a substitute for reading; his July 2026 “gauntlet” post, the June 2026 pipeline description, and the swarm-forge repository give the full picture.
- Faros AI, “The AI Productivity Paradox Report” (2025) and “The Acceleration Whiplash” (2026) — the telemetry behind the review-time, unreviewed-merge, and incident figures.
- Yu et al., “Habituation at the Gate” (2026) — approval rising and scrutiny falling among repeat reviewers of agent PRs.
- Duma et al., “These Aren’t the Reviews You’re Looking For” (2026) — most AI-generated PRs receive no human review.
- Google engineering practices, “What to look for in a code review” and Software Engineering at Google, chapter 9 — the canonical list of what a reviewer is meant to do and why.
- Swarmia, “Should humans still review all your code?” — a risk-tier proposal and the observation that review rarely catches what breaks production.
- Petrović and Ivanković, “State of Mutation Testing at Google” (2018) and Petrović et al., “Practical Mutation Testing at Scale” (2021) — diff-based mutation inside mandatory review, and why coverage alone misleads.
- Birgitta Böckeler, “TDD inside the agent loop” and “Maintainability sensors for coding agents” — why agent-written tests can confirm the implementation instead of the spec, and what computational sensors do and do not fix.
- Gergely Orosz, “TDD, AI agents and coding with Kent Beck” — agents deleting failing tests, and TDD as the counter.
- Kieran Klaassen, “I Stopped Reading Code. My Code Reviews Got Better.” — thirteen parallel review agents on a 27-file change.
- Anthropic, “Bringing Code Review to Claude Code” — parallel find, verify, and rank; the 16% to 54% coverage figure.
- Cursor, “Closing the code review loop with Bugbot Autofix” — bug resolution at merge rising from 52% to 76% once the reviewer could also propose fixes.
- Greptile on Martian’s Code Review Bench and CodeRabbit on the same benchmark — the F1 and recall range for current AI reviewers.
- Thornton, “Can Adversarial Code Comments Fool AI Security Reviewers” (2026) — detection rates on a vulnerability benchmark, and the gain from cross-checking static analysis.
- Macroscope, “The End of Human Code Review” and “What Is Approvability?” — the always-on review thesis and a concrete auto-approve and always-escalate policy.
- Software Engineering at Google, chapter 22, “Large-Scale Changes” — Rosie and pattern-based auto-approval as the pre-LLM precedent.
- GitHub Docs, “Automating Dependabot with GitHub Actions” and Will Larson, “Automatically merging dependabot PRs” — the patch-version auto-merge pattern and its CI precondition.
- GitHub changelog, “Copilot code review can now approve pull requests” (1 September 2026) — a non-human approval counting toward required approvals.
- Cory House, “Don’t review code. Review decisions.” and his five-second check script — the decision-level framing and gate-readiness inside the agent loop.
- Addy Osmani, “Agentic Code Review” — the human moves up a level; risk-tiered depth.
- Itamar Friedman on Tech Lead Journal, “The Future of Code Review” — reviewing rules, workflows, and agent transcripts instead of lines.
- Blake Crosley, “Agents Supersede the Reviewer, Not the Review” — review moving to specification and accountability.
- Dex Horthy, reply to Uncle Bob (August 2026) — the sharpest short statement of the case for still reading the code.
- OpenAI, “Review GitHub pull requests with Codex” — review rules do not replace tests, branch protections, or required approvals.
- Georgia Tech, “Bad Vibes: AI-Generated Code is Vulnerable” and Rabbi et al., “Insights into Security-Related AI-Generated Pull Requests” — what ships when there is no gate.
- DORA, 2025 State of AI-assisted Software Development — throughput up, stability down, and the control-systems explanation.
- Mitchell, Ghosh, and Passi, “AI Agents Push Humans Out of the Loop” (2026) — the skill-erosion argument for keeping people engaged with the system, if not with every diff.
Footnotes
-
Response on X, 14 April 2026, to a query from the Wookash Podcast about a 500-line Claude Code PR review. The July comment comes from a response on 23 July 2026 that received several million views. Both are included in the sources. ↩
-
Faros AI, “The AI Productivity Paradox Report”, July 2025, using telemetry data from 10,000+ developers in 1,255 teams: 98% increase in pull requests merged, 91% increase in review duration, 154% increase in pull request size. These figures compare teams at different stages of AI adoption, so they represent associations, not results from a randomised trial. ↩
-
Faros AI, “AI Engineering Report 2026: The Acceleration Whiplash”, April 2026, looked at 22,000 developers and more than 4,000 teams across two years. Median time in review rose by 441.5% (mean up 199.6%). Pull requests merged without review increased by 31.3%. Incidents per pull request went up by 242.7%. The 441.5% number is frequently cited without reference to the source. It comes from the median in this 2026 report. ↩
-
Yu et al., “Habituation at the Gate”, arXiv, June 2026, on 11,429 reviews from the AIDev dataset: approval rate increased from 30.1% to 36.8%, inline comments decreased by 22%, response time increased by 3.5 times. The sample consists of open-source projects, where reviewer time is more limited than on most paid teams, so the size of the effect may vary. The direction is difficult to dispute. ↩
-
Duma et al., “These Aren’t the Reviews You’re Looking For”, arXiv, May 2026. ↩
-
Martin’s own limits, drawn from his X posts: cyclomatic complexity under approximately 4 per function, and a CRAP score of 6 or less. CRAP merges complexity with test coverage, so complex code remains affordable only when it is thoroughly tested. ↩
-
Petrović and Ivanković, “State of Mutation Testing at Google”, ICSE 2018, noted 6,000 engineers using the system for every change they wrote or reviewed. The 24,000-developer number comes from the 2021 follow-up, “Practical Mutation Testing at Scale”, by Petrović, Ivanković, Fraser, and Just. These two are often confused. ↩
-
Gherkin is the plain language “Given, When, Then” format for describing acceptance scenarios that Cucumber-style tools use. Martin uses an agent to convert each task into Gherkin, creates executable tests from it, and then modifies the scenario values to verify the tests actually rely on them. ↩
-
Anthropic, “Bringing Code Review to Claude Code”, March 2026. Both numbers are drawn from Anthropic’s internal use, meaning they reflect a team that already has a solid test harness and regularly works with code written by agents. Anthropic’s product does not approve pull requests; the layering described in this post is what transforms findings into a decision to merge. ↩
-
Greptile achieved 60.8% F1 in July 2026 and CodeRabbit reached 51.2% in March 2026 on the same Martian Code Review Bench. The top performer shifted between these two results while the overall range remained almost unchanged. This benchmark assesses whether comments are adopted, not actual bugs, so consider the recall as the maximum possible fraction of relevant issues that are identified. ↩
-
Thornton, arXiv 2602.16741, February 2026: eight models, 9,366 trials, conducted on a vulnerability set of 100 samples. The set is small and the flaws are inserted, therefore real-world recall will be lower. The ordering is the useful part. ↩
-
In the “Large-Scale Changes” chapter of Software Engineering at Google by Hyrum Wright, this approach is described. Each change is atomic within a single project and the approval tool operates based on patterns. This provides a precedent for automated approval of a category of change, rather than for individual judgment. ↩
-
GitHub changelog, 1 September 2026. The feature starts disabled and the approval is removed when new commits arrive. This matches the correct structure for an initial policy: enable manually, apply to each repository, cancel on modification. ↩
-
From his X post dated 1 June 2026 and the README file for his swarm-forge repository. The six-agent version he executed on 4 June 2026 included a specifier, coder, cleaner, architect, hardener and QA agent. ↩
-
In April 2026, Georgia Tech’s SSLab scanned over 43,000 advisories. Also in April 2026, Rabbi et al. published arXiv 2604.19965. Both studies examine AI-written code in general, not code that went through a layered gate. That is the key point: the failures arise from pipelines that lacked such a gate. ↩
-
Mitchell, Ghosh, and Passi, “AI Agents Push Humans Out of the Loop”, arXiv, August 2026. A randomised trial conducted at Anthropic, mentioned in the working-with-agents post, showed the same result for individuals: engineers who delegated code generation without raising questions had the least understanding of the outcome. ↩
