Don’t Scale the Reviewers. Engineer the Agent’s Environment.
Beyond Coding — Florian Buetow, Xebia, with Patrick Akil — “AI Architect: Why The Best Engineers Are Solving Code Review Bottlenecks,” W1uG25of2t0
Generate ten times the code and the bottleneck is no longer typing. It is senior engineers reading it, burning out, and accumulating cognitive debt they cannot pay down. Florian Buetow, an AI engineer at Xebia, had just come back from Google I/O when he sat down with Patrick Akil. Even Google, he said, would admit code review is the bottleneck and they do not know how to solve it yet. One serious answer is: stop doing code reviews. The next question is how you dare.
His answer is not a bigger queue of Copilot comments on every PR. It is an environment that gives the agent the feedback a reviewer would have given, as close as possible to the laptop where the code is generated — not after the commit, not in GitHub. Formatters and SonarQube already existed. Humans used to take that output and fix the file. Agents can take it themselves if you wire it in, in natural language, as a stop hook. That is the whole move: stop being the human in the loop for common mistakes.
Horizontal automation versus a custom environment
Big companies publish crazy efficiency numbers — Spotify automated an entire deployment layer, Google reported 50% of code as AI-written in 2025 and was pushing toward 75% — and almost never the translation manual. Buetow’s read of what they actually publish is that review is becoming policy, not disappearing. Amazon, after outages and revenue loss he attributes to AI-generated code, required senior review on some systems before merge or deploy. Scrutiny is now stratified. Those anecdotes are his, from public fragments, not a dump of internal postmortems.
He splits the industry into two axes. Horizontal scaling automates the pipeline you already have: every PR gets an automatic Copilot review. Companies do this and then stay quiet about quality. Vertical scaling is a small team building custom tooling so the product ships the way they intend — a purpose-built environment for their agents, not a blueprint dropped on every repo. Stack: model, then harness around the model, then the environment the harness runs in.
Does the harness matter — Copilot, Claude Code, Codex? Immensely. In his experience the harness matters more than the model. It supplies tools, prompting, memory, and the ability to execute a tool call. He tried to implement a tool from a full specification suite. Spec-driven development: specify exactly enough and the model will implement exactly. “Unfortunately not true.” It failed in the usual way: perfect prompt, model does something else. Then he generated the parallel tests up front, TDD-style, and let failing tests be the feedback. Same frontier model, two harnesses: it worked in one and not the other. At the time the winner was Claude Code. Later he would have said Codex for implementation. Moving target. Do not freeze “we must only use Claude Code” into policy. Models have different personalities — instruction-following versus filling gaps — and the next release can invert last quarter’s winner. You cannot stop experimenting.
Stop hooks, Ralph loops, Semgrep
The feedback cycle he means is automated. On CLI tools, a stop hook fires when the agent thinks it is done. Wire it to a shell script that runs the suite and the guardrails. Engineer those guardrails to emit natural language: this is forbidden, do it this way — the prompt you would have typed as a human. Pair that with a Ralph loop (Codex and Claude now have a functionally similar “keep going until it is fixed” command) so the model eats the error and continues. Goal and Ralph, he said, are the same idea. He does not claim to know the exact implementation.
The guardrail that surprised him most is Semgrep. Regex over code constructs. His standing example: no default values on Python method parameters. In his experience that is one of the greatest sources of later review and debug pain. Policy: you must not write it that way. Whenever he interrogates the model about code instead of reading it, and the explanation is nonsense, that becomes another rule. Over time the environment tightens toward his preferences — not only taste. Code is context for the next agent. Vibe-coded mess confuses the model later. Sustainability did not change: simple and easy to change. What changed is we allow ourselves more volume of code that is neither, because it works, until the next layer is brittle.
Modularity still helps. Clear boundaries, interfaces that are not allowed to change. That is the other high-value class: architectural unit tests — cheap, fast, they only inspect dependencies. Forbid the UI from talking to the database; it must go through business logic. Left alone, models draw interconnection graphs a human would never draw. Encode the rage into another architecture test. That is a guardrail too.
What still cannot be a hook
What remains human: what we want to build, how the system is structured, how we keep it maintainable. Models are not there yet on architecture. His workload is now: understand exactly what to build so there is no doubt, then sketch services, modules, functions — the entire architecture except implementation — and encode that as rules. Cognitive debt, in his telling, arrives when you no longer understand how components talk. Combat that or you cannot reason about the system at all. Patrick’s version of the remainder: implementation is already being automated; with enough guardrails, review might collapse to essence — behavior, maybe the spec — while understanding the system stays the engineer’s skill.
Patrick’s seat in a large org: you may not already know the system, so the investigation that used to happen “as you go” now has to happen up front. Being messy is newly expensive. Some people resist because all the hard work moved to the start. Buetow’s retort: how did you ever develop software if you did not know what you wanted to build? Same work, earlier, more intense. That becomes a discipline. For juniors who can no longer “learn by typing the code,” this is the skill: the pencil-and-paper engineer staring at a screen before hands-on, versus the one who has to start typing to think. Prototyping with AI before the spec is stone is, he said, incredibly rewarding — an hour talking to a model exploring a tangent, then thinking at product level instead of “how do we write this in code.”
The cost is real. Parallelized work exhausts people. A developer last Wednesday, brains on standby, constantly reacting because the model takes twenty minutes. Combat is discipline: stay in one project context. While waiting, do not hard-switch products. Iterate on the environment itself — that is its own project — or start another session on the same codebase and interrogate it. If you do not, you type “what did we do in the last half hour?” Claude Code now auto-summarizes a stale session. Ownership still distinguishes people who care about the craft. Patrick cited a conversation about cognitive surrender: the agent takes the wheel; if it fails it is the agent’s fault; if it works, also the agent. Employees, Buetow said, hand people a grenade and say do not blow it up, but use it. Amazon-style policy — do not YOLO the billing system — will be normal. YOLO elsewhere, maybe.
Review the spec, generate the small test
“No code reviews” is a loaded phrase. Attempt it and you get cheap deterministic wins first, then the conversation moves to architecture and validating specifications before code. You can still use a specialized agent to scan for architectural violations and treat that as the starting point for a human look. Minimize human-in-the-loop because generation is 10× or 100×; everything downstream is crushed. There is no remaining excuse not to write a test for a bug — generate it from the behavior. A small generated test fails less often than “generate me a microservice.”
Patrick has spoken to teams that fully adopted spec-driven development, review the spec rather than the diff, and seem happy; he wants them on the show. Buetow likes specs mostly because they give humans clarity. Give the same spec to a model and in five minutes it has drifted because something was underspecified. Iterate the document toward perfect, switch models, it breaks. Specs should stay as a shared-understanding document, not as if they were code. Fine-grained behavioral specification matters more. TDD’s old claim — delete the software, keep the tests, rebuild — he saw work for the first time when behavioral tests fed the agent and a spec started the run. Existing large TDD codebases? Burn some tokens over the weekend.
The playing field flattened. Harnesses change in months; big-company blog posts are less of a north star. His education method is projects at different scales: one never-looked-at vibe-coded toy; one TDD-plus-behavioral-tests-plus-guardrails; one multi-microservice system where he adds features and watches what the model does wrong, then dumps that into a default-guardrail repo per language. Ask the model to restate its understanding of the task. When subagents arrived he had zero visibility into what they told each other, so he spawned them in another terminal. First-step handoffs already deviate. Looking under the hood taught him more than the tools’ marketing.
This week
Step one on a live production codebase: guardrails. Formatter, linter, Semgrep rules for anti-patterns you can ask the model to list. Team resistance? Run it locally on files you authored. Measure, even impressionistically, with versus without. Once you do more work with less babysitting, you will not go back. Capture PR feedback from teammates as more rules. Data-mine session logs under ~/.claude: “where did I repeatedly remind you?” Turn that into a static check. He said writing that as a skill takes about fifteen minutes.
Stuck with one corporate harness? Organizations want to standardize for five years; that is not how this works. If you cannot choose, find what that harness is actually good at — PR docs, debugging — and use it there. If you have one experiment’s worth of time: encode one piece of human PR feedback as a Semgrep rule. No default parameter values. Never swallow errors; always propagate. Custom for this project. You will see wins quickly even though “we already had static checks.” You did not already encode your taste as the agent’s environment.