skip to content
Karisse's Dev Blog

Uncovering AI's Failure Modes From Code Generation to Code Review

I dug into a month of my Claude development history to understand what AI gets wrong. Perhaps the path to autonomous development doesn’t start with giving AI more autonomy, instead it should start with getting much better at defining the boundaries of that autonomy.

7 min read
cover image

Image generated with ChatGPT

Introduction

After looking at a month of my own development history with Claude, I found something that made me much less confident about automating the PR approval step.

I recently read The Pragmatic Engineer’s article on code reviews where the author discusses about how there are companies who have replaced human code reviews with automated AI reviews for production entirely. I’m still firmly in the four-eyes camp: human review, with an addition of an AI reviewer. But reading how there are companies that could automate code reviews with AI inspired me to challenge my perspective. Are AI reviews truly reliable enough such that we can automate code reviews with AI?

I started this investigation asking whether AI could replace human code review. But the data gave me a different perspective: before asking whether AI can review code reliably, we need to understand what kinds of failures AI is actually introducing during development, and whether those failures can actually be addressed.

Signal or Noise?

I have done some failure discovery which I will be sharing in this article but before diving into the topic deeper, let’s first set some context. There are a few narratives circulating in the industry that set the context for this discussion based on the articles that I’ve read and people whom I’ve spoken to.

image.png
Image Generated with ChatGPT

The general sentiment seems to be:

  1. Majority of the code generated are now generated by AI, leading to higher code velocity and more PRs to review
  2. PRs have gotten bigger and bloated where multiple changes are now done in a single PR which increases blast radius and time to review
  3. Shipping velocity hasn’t improved much as human approvals are still needed

And somehow, the solution to these pain points is to automate AI reviews. But how reliable is this solution? Is this the only solution? Can we refine other facets of the development process?

Uncovering What’s Beneath The Sheets

Rather than speculating whether AI reviews are reliable, I dug deeper to look at what is actually happening in my own development workflow. I got Claude to pull out my past Claude sessions from the past 1 month to analyse what was caught during human review at the pre-PR stage, the quality of AI reviews during PR stage, and the current gaps in the existing AI development process.

The goal here is to perform failure discovery on actual data to understand where the gaps are, uncover the mistakes and failure modes made by AI that were caught during human review, and compare this with the quality of AI reviews when a PR is raised.

Failure Discovery

The below table showcases the severity of the mistakes caught during pre-PR stage during a human review.

SeverityDefinitionCount%
CriticalWould have shipped a functional bug, broken acceptance critera (AC), a security/PII issue, or a data-correctness issue1820.2%
ModerateReal defect or gap, not production-breaking5255.3%
MinorStyle/preference, no functional consequence2425.5%

Reviews performed by AI, specifically Claude and Gemini, are done at the PR stage. In a retrospective analysis using the review instructions Claude and Gemini had, the AI reviewers would likely have caught only 22% of the critical issues that humans caught during pre-PR review. This is a small dataset from my own development workflow, so I’m treating it as failure discovery rather than a benchmark of AI code review accuracy.

The next table is more interesting where it explains why the mistakes happened and Claude’s failure modes.

CategoryDescriptionCountCritical / Moderate / MinorWhy it happens
Correctness / logic bugWrong behavior, wrong output, or a missed edge case1710 / 7 / 0Claude reasons about a condition locally against the diff instead of enumerating every state it partitions — especially for shared/proxy conditions used in more than one place
Architecture / design disagreementBoth approaches work, but redirected to a different design/pattern182 / 12 / 4Claude solves a problem the first reasonable way it thinks of, without first checking whether the codebase already has an established pattern for that exact class of problem
Wrong assumption / hallucinationAsserts something false about the codebase/data and builds on it82 / 6 / 0A gap in Claude’s knowledge of current code/data state gets filled with a plausible-sounding default instead of being verified by reading the actual source
Requirements / AC mismatchCode works but doesn’t satisfy what the ticket needed72 / 3 / 2Claude fixes the specific instance it’s shown without checking whether the same requirement applies to sibling cases (other queues, other test scenarios)
Repo convention violationBreaks a documented project standard150 / 11 / 4Most of these conventions aren’t written down anywhere (see detail below); the ones that are documented are worded too generally to prevent the specific failure
Over-engineering / scope creepMore than the ticket asked for; trimmed back160 / 10 / 6Claude defaults to defensive/comprehensive implementations (extra guards, validation, adjacent investigation) rather than the minimal change the ticket actually needs
Test coverage gapsMissing cases or non-compliant mocking50 / 2 / 3Tests are written to mirror the implementation Claude just wrote rather than derived independently from the spec, so implementation gaps propagate into the tests
Style / naming / nitPure preference, no functional consequence50 / 0 / 5Naming/formatting decided in the moment without checking conventions already used in sibling code
Security / PII / data-safetyExposed data, spoofable identity fields, or similar22 / 0 / 0Security-relevant defaults (trusting a client-supplied field, escaping user input) aren’t the first instinct unless the ticket explicitly calls for adversarial thinking
Wrong root-cause diagnosisWrong underlying cause in a debugging session10 / 1 / 0In a debugging task, Claude anchors on the first plausible explanation matching symptoms instead of ruling out alternatives before proposing a fix

The common theme I see is that Claude can derail from the initial user acceptance criteria provided, perhaps it has “forgotten” about the product requirement document (PRD) and ticket specs after focusing on the implementation detail.

Claude also has a lack of familiarity and understanding of the codebase, where it also does not perform much exploration before implementing. The irony here is that if you have someone new who just onboarded onto your codebase, you’d expect the human to get familiar with it overtime and make less of those mistakes as we naturally build up those context overtime. But with Claude, there’s a lack of such system in place to create a durable contextual understanding of the codebase and domain.

Clearly, a lot can go wrong during development and there’s so much more that can be done to steer the AI in the correct direction. Perhaps improving rulesets and AGENTS.md? Or maybe we need to define checkpoints for review.

def() def() def()

Automating reviews with AI is only part of the solution that everyone is buzzing about, one step that I think we should not be skipping is defining; defining based on past data what can or cannot be automated with AI reviews and rulesets reliably is important too.

Define what good looks like. Define what bad looks like. Use past failures to define the boundary between what AI can safely automate and where humans still need to be involved.

The Automation Gap Is Larger Than It Looks

Only 22% of the critical bugs would’ve likely been caught by AI reviews in the past 1 month. That’s a pretty big gap if the goal is a fully automated PR approval. If anyone is still holding a high bar of what quality and reliability looks like, I believe we’ve got a mountain to climb to be able to achieve something that is automatable at the same time.

Before we automate the final approval step, we need to understand the failure modes, define what good and bad look like, and measure where AI consistently succeeds and fails. Perhaps the path to autonomous development doesn’t start with giving AI more autonomy, instead it should start with getting much better at defining the boundaries of that autonomy.