When AI Reviews and Nobody Reads, You Configured the Wrong Thing

Ankit Jain, Aviator — AI Engineer — “How to Kill the Code Review,” YgEv7IQzGdM


Title slide: How to Kill the Code Review, by Ankit Jain of Aviator. Tagline: The math broke. The gate broke. The theater’s over.
Opening slide: “How to Kill the Code Review. The math broke. The gate broke. The theater’s over.” Ankit Jain, Aviator, at AI Engineer World’s Fair. ▶ 0:25

Ankit Jain is not asking when teams will stop reading diffs line by line. He says they already have. Volume is up. Incidents per PR are up. Median review wait is four times what it was, because coding got “solved” and everything stuck at review. Over 30% of changes merge with no review at all. He cited 861% code churn as the production side of that same squeeze. Those figures are speaker-reported, the same acceleration story circulating this conference. His sharper claim is about the replacement ritual, not the missing humans. The slide behind him puts four numbers in a row, sourced to a Faros AI “Acceleration Whiplash” report from April 2026 covering 22,000 developers across 4,000 teams: +861% code churn (deleted vs. added), +243% incidents-to-PR ratio, +441% median time in review, and +31.3% of PRs merging with no review at all.

Slide titled We have already stopped reviewing it, with four stat cards: +861% code churn, +243% incidents-to-PR ratio, +441% median time in review, +31.3% of PRs merge with no review at all
“We’ve already stopped reviewing it.” The four stat cards, sourced on-slide to Faros AI’s April 2026 acceleration report (22,000 developers, 4,000 teams). ▶ 1:58

An AI writes the change. Two or three AI reviewers argue with it in the GitHub UI. The thread ping-pongs until the comments resolve. A person skims and merges. Jain’s line: when AI reviews and nobody reads, we have configured the wrong thing.

Diagram: Agent writes and AI reviews and revises loop (iterates N times, no human in the loop), then human skims, then merge
“AI writes the code. AI reviews the code. Why is there a UI?” The agent and the AI reviewers iterate with no human in the loop; a person skims and merges. Footer: when AI reviews AI in a UI nobody reads, you’ve configured the wrong thing. ▶ 2:45

He is co-founder of Aviator, which is building an AI code-verification platform, and he is here to sell a product called Verify. He is also here to correct a LinkedIn / Latent Space post from a few months earlier: a five-layer trust model for merging without line-by-line review. He got some of it right. He got the purpose of review wrong. The post’s model, shown on a slide with a QR code to latent.space/p/reviews-dead, is a cheese-slice diagram of five layers: compare multiple options, deterministic guardrails, acceptance criteria, permission systems, and adversarial verification, with the caption “building trust through layers.”

Slide: The 5-layers of trust, showing a Swiss-cheese diagram with five slices: compare multiple options, deterministic guardrails, acceptance criteria, permission systems, adversarial verification, and a QR code to latent.space/p/reviews-dead
The earlier post: “The 5-layers of trust.” Specs and agents go in on the left; five layers of imperfect filters stand between them and a merge. ▶ 0:50

Review was never only about bugs

Formal code review is young. Google’s Mondrian made it a thing internally in 2006. Early Windows, he noted, shipped without it. Catching bugs, conventions, and security issues is the part everyone names. The part his five-layer model missed is alignment: knowledge sharing, mentorship, architectural feedback, onboarding, collaboration. If you are vibe-coding a solo project, this talk is not for you. If you work on a team and you are not living in a dark factory where nobody looks at code, alignment is the part of review that has to survive. Semantic accuracy can be tooled. Alignment cannot be discarded.

Slide: Code review is not about code review. Two groups: semantic accuracy (catch bugs, conventions, security, spot regressions, naming, edge cases) and alignment (knowledge sharing, mentorship, architecture feedback, onboarding, shared ownership, API design)
“Code review is not about code review.” Semantic accuracy (bugs, conventions, security, regressions, naming, edge cases) on the left; alignment (knowledge sharing, mentorship, architecture feedback, onboarding, shared ownership, API design) on the right. Footer: “Semantic accuracy needs better tools. But alignment must survive!” ▶ 4:15

Spec-driven development looks like the adult replacement: write a complete spec, hand it to an agent, verify. Jain’s objection is that this is the 1970 waterfall model — requirements, specification, implement, verify — with no feedback loop. The spec is written before you know what you will learn. That is why people still sit in Claude Code, Codex, Cursor sessions. Things were not clear. As you implement, you find more issues, and you do not go back and update the spec because the spec is “done” and you expected deterministic code. LLMs are not deterministic. They make decisions. Spec-driven work is a useful methodology that falls short of day-to-day software. What should be carried forward is intent.

Slide: Spec-driven development is waterfall in new clothes. Waterfall 1970: requirements, design, implementation, verification equals Spec-driven 2026: spec, architecture, generation, verification. Bullets: no feedback loop; specs drift
“Spec-driven development is waterfall in new clothes.” Waterfall (1970) on the left, spec-driven (2026) on the right, with the same four-step shape. Below: no feedback loop, and specs drift, because decisions made during implementation never make it back into the spec. ▶ 5:45

Intent does not live only in the spec. It lives in the Jira ticket (the goal), in PRDs (the plan), and most importantly in the prompts. That is where the real decisions happen: back and forth with the agent, starting from a ticket. Then the team opens a pull request and throws the prompts away. That is the thing he wants changed.

Slide: Intent does not only live in a spec. Three cards: JIRA ticket (the goal), PRDs and design docs (the plan), and your prompts, highlighted, where the real decisions get made
“Intent doesn’t only live in a spec.” The Jira ticket is the goal, the PRD is the plan, and the prompts, highlighted, are where the real decisions get made, then thrown away when the session ends. ▶ 6:25

An AI slop registry, then a different surface

Alignment does not excuse bugs. LLMs are not great at catching them, and AI reviewers are not perfect. If you are still doing any manual review — and he expects some degree of it — you are probably writing the same comments over and over. Capture those and you get what he calls an AI slop registry: recurring comments become guardrails you do not have to leave again. Do it a few times and the system learns from human review experience on top of the base model. Every merge can compound the registry instead of compounding the comment load.

Slide: The AI Slop Register. A Verify invariants screen listing active rules such as use the Money type for currency amounts, use the structured logger not print, no direct writes to user tables, error paths increment a metrics counter, no hardcoded secrets. Side panels: codify the comment, and it builds itself
“Codify the comment: the AI Slop Register.” A review comment (“use Money, not float”) becomes an invariant that rejects float fields on currency. The register’s Review tab fills weekly from merged-PR comments, and a human approves each rule before it goes live. ▶ 7:50

The two halves close in one loop. Capture the coding session — the questions the agent asked, the answers the human gave, even on a simple task — and those user decisions become acceptance criteria. Criteria plus the slop registry become a test plan. A verification system spins up a preview and runs the plan end to end: even if the code looks right, does it work? That package is the new review surface. Not the diff. Intent: did we implement the capability we defined? Behavior: did it meet the criteria? Architectural argument still happens. It happens one level up.

Slide: Two halves, one loop. Capture: session and acceptance criteria. Build: AI slop register, test plan. Run: preview env, results. Review: reviewer reviews intent, architecture and verdicts, not a diff
“Two halves, one loop.” Capture (session → acceptance criteria), build (AI slop register → test plan), run (preview environment → pass/fail with evidence), and review, where the reviewer looks at intent, architecture and verdicts: not a diff. ▶ 9:08

He walked the Aviator version. Session to criteria, LLM-assisted because writing test plans is painful. Criteria plus invariants to a plan. Verification against previews. For a new feature you might not maintain a frozen suite at all — tests created in real time — while the human’s job is governance: reviewing the plan, not the code. He analogized it to behavior-driven development more than twenty-year-old TDD. The plan is in English. Product managers and designers can participate. Deterministic checks where they can run; LLM as fallback where they cannot. Example: a new payment form. An agent browses the app, fills the form, captures screenshots and database snapshots as evidence. Reviewers look at that evidence and at the session’s intent — what we said we would build, what we tried and rejected — not at every line.

Slide: The session becomes the criteria. A Claude Code session on the left with lines tagged DECISION; on the right an Aviator Verify acceptance criteria panel with four items and an Agree button
“The session becomes the criteria.” Left: a Claude Code session adding a rate limit to an exports endpoint, with the human’s calls tagged as decisions (use the existing RateLimiter util; skip the auth, middleware handles /api/v2; use the Money type, not float). Right: the Acceptance Criteria panel built from them, which the reviewer signs off with “Agree” before any code is judged. ▶ 10:12
Slide: Criteria plus invariants become a test plan that runs. Left, a test plan with runtime and code-scan checks. Right, a run result for PR 1287 with 4 passed, 1 failed: threshold typed as Money failed because it declared float
“Criteria + invariants → a test plan that runs.” The plan mixes one runtime check (fire 101 req/min, expect 429) with code scans. The run against the preview environment passes the 429 and RateLimiter checks but fails “Threshold typed as Money”: it declared a float at exports/service.py:142. That is the register earning its keep. ▶ 11:32

Do not generate the test plan from the code the same agent just wrote. That is the Dexter point from the day before: the writer will not author a plan that catches itself. Session information is the source of truth for criteria. Architecture — data models, how services interact — is what collaboration still needs. Evidence from verification is what confidence needs.

Slide: Deterministic where you can, LLM where you must. Two cards: deterministic check (scan, assertion, status code, screenshots, reproducible) and runtime plus LLM judge (screenshots, fuzzy behavior, for criteria that are genuinely judgment calls). Footer: the model is a tool inside the harness, never the judge of record
“Deterministic where you can. LLM where you must.” Scans, assertions, status codes and screenshots are reproducible; an LLM judge is reserved for criteria that really are judgment calls. “The model is a tool inside the harness — never the judge of record.” ▶ 12:15
Slide: Reviewers review intent, not diffs. Three cards: intent and decisions; architecture; verification results. Footer: the reviewer is more important than ever, they are just not reading diffs anymore
“Reviewers review intent. Not diffs.” What reviewers get instead: intent and decisions (what we set out to build, what we tried and rejected), architecture (the shape of the change, trade-offs worth a human), and verification results (verdicts and evidence). “The reviewer is more important than ever. They’re just not reading diffs anymore.” ▶ 13:58

Homework, a J-curve, and a pilot

Mine your last 1,000 review comments. Build a slop registry for the repeatable ones. A vast majority of comments are repeats. Each capture means you do not have to write it again. Codify semantic accuracy without killing the collaboration half.

Slide: Build the register. Do not kill review on day one. Mine your last 1000 review comments and build an AI slop register. A J-curve chart of payoff over time dips below baseline, bottoms out at about month 2 dip, then climbs above it. Caption: pain is real, payoff is real
“Build the register. Don’t kill review on day one.” The J-curve: payoff dips below baseline in the early months while the register is built, then climbs past it. “Pain is real. Payoff is real.” It compounds: every merged PR proposes new invariants, the next PR catches more, and reviews shrink. ▶ 15:05

It follows a J-curve. The pain is real. The registry takes time before it pays off. That is also the product beat: Aviator is piloting Verify with early design partners, combining alignment capture with slop-registry checks. Remember one thing, he said: code review is not just about code review. It is about getting the alignment.

Closing slide: If you remember one thing, code review is not about code review. Alignment: diff to intent. Semantic accuracy: the AI slop register. aviator.co/verify with a QR code
Closing slide: “If you remember one thing: code review is not about code review.” Alignment means diff → intent; semantic accuracy means the AI Slop Register. Pilot sign-up at aviator.co/verify. ▶ 15:35