Introduction
After looking at a month of my own development history with Claude, I found something that made me much less confident about automating the PR approval step.
I recently read The Pragmatic Engineer’s article on code reviews where the author discusses about how there are companies who have replaced human code reviews with automated AI reviews for production entirely. I’m still firmly in the four-eyes camp: human review, with an addition of an AI reviewer. But reading how there are companies that could automate code reviews with AI inspired me to challenge my perspective. Are AI reviews truly reliable enough such that we can automate code reviews with AI?
I started this investigation asking whether AI could replace human code review. But the data gave me a different perspective: before asking whether AI can review code reliably, we need to understand what kinds of failures AI is actually introducing during development, and whether those failures can actually be addressed.
Signal or Noise?
I have done some failure discovery which I will be sharing in this article but before diving into the topic deeper, let’s first set some context. There are a few narratives circulating in the industry that set the context for this discussion based on the articles that I’ve read and people whom I’ve spoken to.
The general sentiment seems to be:
- Majority of the code generated are now generated by AI, leading to higher code velocity and more PRs to review
- PRs have gotten bigger and bloated where multiple changes are now done in a single PR which increases blast radius and time to review
- Shipping velocity hasn’t improved much as human approvals are still needed
And somehow, the solution to these pain points is to automate AI reviews. But how reliable is this solution? Is this the only solution? Can we refine other facets of the development process?
Uncovering What’s Beneath The Sheets
Rather than speculating whether AI reviews are reliable, I dug deeper to look at what is actually happening in my own development workflow. I got Claude to pull out my past Claude sessions from the past 1 month to analyse what was caught during human review at the pre-PR stage, the quality of AI reviews during PR stage, and the current gaps in the existing AI development process.
The goal here is to perform failure discovery on actual data to understand where the gaps are, uncover the mistakes and failure modes made by AI that were caught during human review, and compare this with the quality of AI reviews when a PR is raised.
Failure Discovery
The below table showcases the severity of the mistakes caught during pre-PR stage during a human review.
| Severity | Definition | Count | % |
|---|---|---|---|
| Critical | Would have shipped a functional bug, broken acceptance critera (AC), a security/PII issue, or a data-correctness issue | 18 | 20.2% |
| Moderate | Real defect or gap, not production-breaking | 52 | 55.3% |
| Minor | Style/preference, no functional consequence | 24 | 25.5% |
Reviews performed by AI, specifically Claude and Gemini, are done at the PR stage. In a retrospective analysis using the review instructions Claude and Gemini had, the AI reviewers would likely have caught only 22% of the critical issues that humans caught during pre-PR review. This is a small dataset from my own development workflow, so I’m treating it as failure discovery rather than a benchmark of AI code review accuracy.
The next table is more interesting where it explains why the mistakes happened and Claude’s failure modes.
| Category | Description | Count | Critical / Moderate / Minor | Why it happens |
|---|---|---|---|---|
| Correctness / logic bug | Wrong behavior, wrong output, or a missed edge case | 17 | 10 / 7 / 0 | Claude reasons about a condition locally against the diff instead of enumerating every state it partitions — especially for shared/proxy conditions used in more than one place |
| Architecture / design disagreement | Both approaches work, but redirected to a different design/pattern | 18 | 2 / 12 / 4 | Claude solves a problem the first reasonable way it thinks of, without first checking whether the codebase already has an established pattern for that exact class of problem |
| Wrong assumption / hallucination | Asserts something false about the codebase/data and builds on it | 8 | 2 / 6 / 0 | A gap in Claude’s knowledge of current code/data state gets filled with a plausible-sounding default instead of being verified by reading the actual source |
| Requirements / AC mismatch | Code works but doesn’t satisfy what the ticket needed | 7 | 2 / 3 / 2 | Claude fixes the specific instance it’s shown without checking whether the same requirement applies to sibling cases (other queues, other test scenarios) |
| Repo convention violation | Breaks a documented project standard | 15 | 0 / 11 / 4 | Most of these conventions aren’t written down anywhere (see detail below); the ones that are documented are worded too generally to prevent the specific failure |
| Over-engineering / scope creep | More than the ticket asked for; trimmed back | 16 | 0 / 10 / 6 | Claude defaults to defensive/comprehensive implementations (extra guards, validation, adjacent investigation) rather than the minimal change the ticket actually needs |
| Test coverage gaps | Missing cases or non-compliant mocking | 5 | 0 / 2 / 3 | Tests are written to mirror the implementation Claude just wrote rather than derived independently from the spec, so implementation gaps propagate into the tests |
| Style / naming / nit | Pure preference, no functional consequence | 5 | 0 / 0 / 5 | Naming/formatting decided in the moment without checking conventions already used in sibling code |
| Security / PII / data-safety | Exposed data, spoofable identity fields, or similar | 2 | 2 / 0 / 0 | Security-relevant defaults (trusting a client-supplied field, escaping user input) aren’t the first instinct unless the ticket explicitly calls for adversarial thinking |
| Wrong root-cause diagnosis | Wrong underlying cause in a debugging session | 1 | 0 / 1 / 0 | In a debugging task, Claude anchors on the first plausible explanation matching symptoms instead of ruling out alternatives before proposing a fix |
The common theme I see is that Claude can derail from the initial user acceptance criteria provided, perhaps it has “forgotten” about the product requirement document (PRD) and ticket specs after focusing on the implementation detail.
Claude also has a lack of familiarity and understanding of the codebase, where it also does not perform much exploration before implementing. The irony here is that if you have someone new who just onboarded onto your codebase, you’d expect the human to get familiar with it overtime and make less of those mistakes as we naturally build up those context overtime. But with Claude, there’s a lack of such system in place to create a durable contextual understanding of the codebase and domain.
Clearly, a lot can go wrong during development and there’s so much more that can be done to steer the AI in the correct direction. Perhaps improving rulesets and AGENTS.md? Or maybe we need to define checkpoints for review.
def() def() def()
Automating reviews with AI is only part of the solution that everyone is buzzing about, one step that I think we should not be skipping is defining; defining based on past data what can or cannot be automated with AI reviews and rulesets reliably is important too.
Define what good looks like. Define what bad looks like. Use past failures to define the boundary between what AI can safely automate and where humans still need to be involved.
The Automation Gap Is Larger Than It Looks
Only 22% of the critical bugs would’ve likely been caught by AI reviews in the past 1 month. That’s a pretty big gap if the goal is a fully automated PR approval. If anyone is still holding a high bar of what quality and reliability looks like, I believe we’ve got a mountain to climb to be able to achieve something that is automatable at the same time.
Before we automate the final approval step, we need to understand the failure modes, define what good and bad look like, and measure where AI consistently succeeds and fails. Perhaps the path to autonomous development doesn’t start with giving AI more autonomy, instead it should start with getting much better at defining the boundaries of that autonomy.