The Implementer Is Overloaded. Review Should Commit.

Matt Pocock, AI Hero — AI Engineer Paris 2026, Coding Agents & Software Factories — “Fixing the PR Bottleneck,” LlgiOCmFG_w


The pull request was already the pile nobody wanted to look at. Agents made it trivial to open another one. Matt Pocock’s talk is not about writing them faster. It is about brakes. A software factory, as he uses the word, is what happens when initiation is no longer only human: a classifier like Jev turns an event into a fix or a reproduction, PlanetScale slow-query reports fire another loop, deterministic code keeps shoving work in. Permanent acceleration without brakes is a slop cannon. Code is the environment the next agent operates in. Bad code begets bad code.

He is a developer educator shipping skills at AI Hero — version 1.3 that week — so the talk is also a product announcement. The argument still has a sharp middle: do not put coding standards on the agent that writes the code, and when the reviewer agent finds something, it should commit, not dump comments for a human to triage.

Three layers, and the checks that lie

Automated checks are the deterministic stuff we have had since roughly the 1950s: lint, tests, typechecking, quality metrics. They cost CPU, not tokens, not human attention. You can layer on more of them than you think, and you are probably not being creative enough. They can still lie. Green CI does not mean ready to merge. Automated review and human review exist to catch those lies.

The first lie is the tautological test. Opus 5, he said, got addicted to them. Constant: X post character limit equals 280. Test: expect X post character limit to be 280. The test reasserts the implementation. Rename the constant and it fails. Change the constant and it fails. It is welded to structure, not behavior.

Worse: a test that two UI sections appear in the right order — video section after content plan — that never renders anything. It reads the module source into memory, finds the strings, and asserts their offset in the file. Rearrange the source and the test dies. It never executed the product.

Then tests that cannot fail. useAudioBoost wraps the DOM AudioContext API, which has complicated error modes. The test stubs those methods with fakes. Those modes never fire. Production will. The model is not trying to cheat. It is following instructions and writing tests tied to structure instead of executing code.

Make checks harder to cheat and the quality bar rises. Pocock’s design answer is John Ousterhout’s deep modules from A Philosophy of Software Design: hide complex behavior behind a small interface. Module A: fat implementation, tiny API. Module B: huge surface, each function does little. Testing at A’s interface produces fewer structure-sensitive tests. The job is to force the agent to use that small interface instead of reaching into internals. He has a skill that, even on a vibe-coded mess, surfaces deepening opportunities as an HTML before/after — reduce duplication, create a deep testable module. A companion codebase-design skill defines a vocabulary so the team is not arguing twenty flavors of DDD: locality (how a small change ripples), leverage (what the caller gets from a simple function on a deep module), seams (“I was using seams before they were cool”). Both locality and leverage, he said, are also good for agents.

Implementation is overloaded. Review is not.

Most people still dump coding standards into the implementer. In one context window that agent has to explore, edit files, and debug against the checks. Add standards on top and it performs worse. Implementation is overloaded. Review is underloaded: it receives a diff, should do a little exploration for wider context, and does not implement or debug. Pile the standards there.

His code-review skill reads coding-standards.md in the repo — customizable, team-owned — and checks the diff against it, in a subagent with its own context budget. Uncomfortable, because everyone wants good code on the first shot. He treats it as two context windows: implement until it works, then review until it is good. Red-green-refactor, split across agents. Do not put those standards in global scope or AGENTS.md, where they drown the implementer and may not even be read. Put them where only the reviewer loads them.

People told him they just use Cursor Bugbot or CodeRabbit. He has tried to write a generic review skill that finds all the bugs and does security. Too general and you drown in false positives. Too specific — “find all the TypeScript stuff” — and Rust people cannot use it. Do not outsource automated review. Build coding standards over time, share them, dump the unread docs into that file.

The other default he wants killed: the review agent comments on the PR. That is more work for the human, who now reads verbose notes and decides what to implement. The reviewer should commit. Fix what it finds. Comment only when it has a real question. The human then reviews a nicer artifact. Stop trying to one-shot good code from the implementer alone.

Human-friendly PRs, then never the same comment twice

After checks and automated review, maximize what a human can do in the remaining minutes. A PR skill is coming into the repo — still in progress when he spoke. Steal the best ideas. First: not every review is essential. Use the AWS one-way-door / two-way-door language. Most PRs are two-way: merge and revert. Civil engineering is one-way; software usually is not. Nuance: a “simple” change that blasts email to 60,000 people is a one-way door. Expensive migrations, data loss — review the hell out of those. Also write the blast radius: what can go wrong, and how bad. He puts a merge-danger summary at the bottom: two-way door, localized radius, pay less attention.

Second: make the why fast. Pseudo-code and diagrams beat paragraphs. He credits the Show Me skill from Dex Hadley’s Human Layer skills repo — less text, more images. Mermaid, UML, or a tiny CLI sketch: new command, two flags. Hard to overstate, he said.

Third: when you human-review, you are not only reviewing the code. You are reviewing the system that created it. Never write the same comment twice. Never catch the agent doing the same thing on two PRs. The mechanism is a new skill, Retro. Feed it a session, a PR plus session, or a week of PRs and reviews. It suggests automated checks, updates to coding-standards.md, navigation pointers in AGENTS.md for places the agent struggled to find, tool-economy fixes (token-waste that is hard to see from outside), and bloat in steering files and skills. Human review then raises the quality of the next human review.

Goal: make human review faster by leaning on the first two layers until it is as optional as you can stand. You do not need to review every two-way door. You do need to review every one-way door. Skills at aihero.dev/skills. He would be in the lobby.