Not Every Line Needs Your Eyes. Every Change Still Needs Proof.

Alex Volkov, ThursdAI — AI Engineer World’s Fair, leadership track — “Should AI Engineers Still Read Code in 2026? The Z/L Continuum,” ZpK5PWX2YRM


Two talks, same conference, opposite standing ovations. Ryan Lopopolo of OpenAI said code is free, deleted his IDE, and told the room humans should stop concerning themselves with implementation. Mario Zechner, creator of Pi, said slow the hell down, agents compound errors with delayed pain, and if the code is critical you read every line. Alex Volkov named the line between them the Z/L continuum — Zechner to Lopopolo — and then admitted he had framed it wrong.

It is not a personality test. It is a routing problem. The question is not whether AI engineers should still read code in 2026. It is what proof this specific change needs. Code got cheap. Attention did not.

Babysitters, not handcrafters

Volkov hosts ThursdAI, a weekly show he has run for three and a half years, and works as an AI evangelist at Weights & Biases and CoreWeave. He has been at every AI Engineer conference since 2023. This one, he said, was three times last year’s size: about 7,000 people, 36 tracks. The Token Maxing room he was speaking in was the wrong place to find people who still hand-type most of their code. One shy holdout, maybe. Everyone else supervises. “We babysit agents.”

The backdrop is the December 2025 break in the trendline. Swyx is collecting receipts at wtfhappened2025.com. METR’s curve, in Volkov’s telling, showed models completing tasks that would take engineers more than 16 hours, then going “way up the curve.” Boris Cherny, creator of Claude Code, now authors 100% of his code through the tool, still ships 20 to 30 PRs, and recently talked about deleting his IDE. Anthropic has said 80% of its code is AI-written — a figure Volkov flagged as already months old and likely higher. GitHub, he said, is on track for 14 billion commits this year against a billion the year before, most of it AI-assisted. Those numbers are speaker-reported, some of them already dated on stage. They are the weather system for the argument, not independent audits.

Same anxiety, two podiums

Lopopolo’s clip: models are good enough that they are isomorphic with writing code. Tools solve real problems in real codebases. Code is free to produce and free to refactor. The important thing is the prompt and the guardrails. You can say “do not produce slop” and refuse slop — if you take short-term velocity hits to double-click into what the agents are struggling with. Volkov’s color: Lopopolo opened like a “talking billionaire,” invented the talking-billionaire lounge in front of the leadership track, and sits in the AGI-pilled corner of OpenAI, where tokens are also free. The Codex dangerous-permissions skip got nicknamed YOLO Popolo. Volkov asked him. He was okay with it.

Zechner’s clip, next day: everything is broken. Teams announcing a product 100% built by agents — congratulations, it sucks. Agents compound “booboos” with zero learning, no bottlenecks, and delayed pain that lands on you. If you do not read the code, you cannot be the person who gets the call when users scream. Non-critical code: write slop ahead. Critical code: read every line.

Those two talks, Volkov said, are the sixth- and seventh-most-watched AI Engineer videos of all time. The hallway already had the Slack argument. He put both men on a line and started asking speakers where they sat. Then he took a show-of-hands in the room. Who has committed code they never looked at? A lot of hands. Who still reads every line, or at least critical code? One cowboy. Volkov wanted to talk to him after.

Output is real. So is the Christmas tree.

Checked against data, the optimists are right about volume. A Faros AI survey from April 2026 — 22,000 engineers, “acceleration whiplash” — is the source he kept citing. Favorite stat: an 861% increase in code deletion per PR. Anthropic, separately, said they are shipping eight times more code per quarter than in 2025.

Then he put up a status page and asked whose it was. Claude. Anthropic is the company that probably uses the most AI-generated code, and the page looked like a Christmas tree. He said he was not there to dunk; scale and other factors exist; GitHub is famously wobbly too. Output does not mean stability. Same Faros study: 31% increase in PRs merged with no review at all, human or agentic. “Don’t do this. I beg of you.” Same study: 242% increase in incidents per PR. A second study: bugs per developer up six times versus 2025. Those percentages are survey- and speaker-reported. They are the bill Zechner said comes due in production.

Anthropic can see the bottleneck. In the recursive-self-improvement essay Volkov cited, they outline a world where acceleration stops and everyone gets used to it, then say that is not likely. What they expect is 10x to 100x to 1,000x output, and then this sentence: as they push more code around the organization, human code review has become a new bottleneck. Amdahl’s law: explode one stage and another blocks. Nobody is removing the human. Anthropic and OpenAI are still hiring. Both still treat human review as a concern.

The mea culpa

The continuum is real. It is not about the people. The same engineer can be Lopopolo on one change and Zechner on another. Different tasks need different proof.

Read closely and the two talks agree more than the meme allows. Lopopolo’s mechanism is moving attention up a layer. Humans are bad at catching the same class of mistake forever. When you catch it in review, write it into documentation, a linter, a reviewer that remembers once so the system catches that type. He is not saying do not inspect. He is saying inspect the system, not every line. Zechner is saying route by task: if it is not critical, let it rip; if it is, read. How do you know what is critical? Zechner’s answer is you read the code. Volkov’s add-on: ask the models. They are good at pointing at the primitive that actually matters in a large repo.

The wrong question is should I still be reading code. The better one: what proof does this specific change need?

Swyx, he said, asked for one screenshot slide. The Monday artifact is a routing table. Authentication, money movement, permissions, irreversible data: you read every line and inspect the critical path yourself. Long PRs glaze the eyes, so split into atomic reviewable PRs — agents are good at that decomposition; ask them. Verification does not go away: traces, evals, shadow mode (he ran out of time; come talk after). Separate the agent that writes from the one that inspects and writes tests. One agent scoring its own exam is not productive. Last: engineer the rails — observability, rollback — because reading spends attention once and engineering makes the system remember. That is Lopopolo’s move, restated as infrastructure.

Capability drift does not retire proof

He coined the continuum 82 days earlier. Then Mythos was announced, Fable landed, only Anthropic had access at first. Jared from Anthropic, on stage: they used to check whether Claude was doing the work right; with Fable they check whether Claude is doing the right work. Volkov said that gave him chills. Andrej Karpathy, newly at Anthropic and taking heat on Twitter, named both ends in one sentence: it has never felt so tempting to stop looking at code at all — but don’t do this in production. Volkov’s joke: this talk could have been that sentence.

On the slide, a capability-drift arrow. The continuum is a temporary place. As models get better, the review layer moves. Yesterday you inspected outputs and read code. Today you inspect task direction and route to the right proof. Tomorrow you may be inspecting loops. Drift changes where proof belongs. It does not remove the requirement.

Loops were the hallway word. Peter Steinberger — OpenClaw, now OpenAI — and Boris Cherny started talking about them within two days of each other. Hands: plenty had heard of loops; fewer were running them. The common fact Volkov wanted in the room: those guys’ tokens are free. Treat them as a lighthouse, Gretzky skating to where the puck is going, not as an instruction to copy the setup on Monday.

His TL;DR: loops are fancy cron jobs. They discover a task, write a prompt from a plan, execute, verify, and try again. An agent that grades itself against the goal with less human intervention. If the builder grades itself, you did not remove review. You hid it. Adi at Google, also at the conference: if he stopped reviewing and relied entirely on automated loops — a Jira bug, the loop picks it up — product quality would suffer, a downward spiral, a deeper hole. Loops do not remove judgment. They raise the stakes on where you put it.

Nobody knows the rest. Anthropic did not know Claude Code would explode into a billion-dollar product. Nobody knew coding agents and harnesses would become the generalized agent, and now OpenAI has Codex, Elon has Grok code, Google has anti-gravity. Capability is jumping. Flexibility is the job. That is why you stay an engineer. Scan the QR for ThursdAI if you want to watch where the line moves — affiliation disclosed, pitch included.

Not every line in 2026 needs your eyes. Every system still needs your judgment.