When AI Reviews and Nobody Reads, You Configured the Wrong Thing
Ankit Jain, Aviator — AI Engineer — “How to Kill the Code Review,” YgEv7IQzGdM
Ankit Jain is not asking when teams will stop reading diffs line by line. He says they already have. Volume is up. Incidents per PR are up. Median review wait is four times what it was, because coding got “solved” and everything stuck at review. Over 30% of changes merge with no review at all. He cited 861% code churn as the production side of that same squeeze. Those figures are speaker-reported, the same acceleration story circulating this conference. His sharper claim is about the replacement ritual, not the missing humans. The slide behind him puts four numbers in a row, sourced to a Faros AI “Acceleration Whiplash” report from April 2026 covering 22,000 developers across 4,000 teams: +861% code churn (deleted vs. added), +243% incidents-to-PR ratio, +441% median time in review, and +31.3% of PRs merging with no review at all.
An AI writes the change. Two or three AI reviewers argue with it in the GitHub UI. The thread ping-pongs until the comments resolve. A person skims and merges. Jain’s line: when AI reviews and nobody reads, we have configured the wrong thing.
He is co-founder of Aviator, which is building an AI code-verification platform, and he is here to sell a product called Verify. He is also here to correct a LinkedIn / Latent Space post from a few months earlier: a five-layer trust model for merging without line-by-line review. He got some of it right. He got the purpose of review wrong. The post’s model, shown on a slide with a QR code to latent.space/p/reviews-dead, is a cheese-slice diagram of five layers: compare multiple options, deterministic guardrails, acceptance criteria, permission systems, and adversarial verification, with the caption “building trust through layers.”
Review was never only about bugs
Formal code review is young. Google’s Mondrian made it a thing internally in 2006. Early Windows, he noted, shipped without it. Catching bugs, conventions, and security issues is the part everyone names. The part his five-layer model missed is alignment: knowledge sharing, mentorship, architectural feedback, onboarding, collaboration. If you are vibe-coding a solo project, this talk is not for you. If you work on a team and you are not living in a dark factory where nobody looks at code, alignment is the part of review that has to survive. Semantic accuracy can be tooled. Alignment cannot be discarded.
Spec-driven development looks like the adult replacement: write a complete spec, hand it to an agent, verify. Jain’s objection is that this is the 1970 waterfall model — requirements, specification, implement, verify — with no feedback loop. The spec is written before you know what you will learn. That is why people still sit in Claude Code, Codex, Cursor sessions. Things were not clear. As you implement, you find more issues, and you do not go back and update the spec because the spec is “done” and you expected deterministic code. LLMs are not deterministic. They make decisions. Spec-driven work is a useful methodology that falls short of day-to-day software. What should be carried forward is intent.
Intent does not live only in the spec. It lives in the Jira ticket (the goal), in PRDs (the plan), and most importantly in the prompts. That is where the real decisions happen: back and forth with the agent, starting from a ticket. Then the team opens a pull request and throws the prompts away. That is the thing he wants changed.
An AI slop registry, then a different surface
Alignment does not excuse bugs. LLMs are not great at catching them, and AI reviewers are not perfect. If you are still doing any manual review — and he expects some degree of it — you are probably writing the same comments over and over. Capture those and you get what he calls an AI slop registry: recurring comments become guardrails you do not have to leave again. Do it a few times and the system learns from human review experience on top of the base model. Every merge can compound the registry instead of compounding the comment load.
The two halves close in one loop. Capture the coding session — the questions the agent asked, the answers the human gave, even on a simple task — and those user decisions become acceptance criteria. Criteria plus the slop registry become a test plan. A verification system spins up a preview and runs the plan end to end: even if the code looks right, does it work? That package is the new review surface. Not the diff. Intent: did we implement the capability we defined? Behavior: did it meet the criteria? Architectural argument still happens. It happens one level up.
He walked the Aviator version. Session to criteria, LLM-assisted because writing test plans is painful. Criteria plus invariants to a plan. Verification against previews. For a new feature you might not maintain a frozen suite at all — tests created in real time — while the human’s job is governance: reviewing the plan, not the code. He analogized it to behavior-driven development more than twenty-year-old TDD. The plan is in English. Product managers and designers can participate. Deterministic checks where they can run; LLM as fallback where they cannot. Example: a new payment form. An agent browses the app, fills the form, captures screenshots and database snapshots as evidence. Reviewers look at that evidence and at the session’s intent — what we said we would build, what we tried and rejected — not at every line.
Do not generate the test plan from the code the same agent just wrote. That is the Dexter point from the day before: the writer will not author a plan that catches itself. Session information is the source of truth for criteria. Architecture — data models, how services interact — is what collaboration still needs. Evidence from verification is what confidence needs.
Homework, a J-curve, and a pilot
Mine your last 1,000 review comments. Build a slop registry for the repeatable ones. A vast majority of comments are repeats. Each capture means you do not have to write it again. Codify semantic accuracy without killing the collaboration half.
It follows a J-curve. The pain is real. The registry takes time before it pays off. That is also the product beat: Aviator is piloting Verify with early design partners, combining alignment capture with slop-registry checks. Remember one thing, he said: code review is not just about code review. It is about getting the alignment.