Engineering · Architecture · Software economics

Cheap Code Does Not Make Every Mistake Cheap

Agents can make a failed experiment disposable. They cannot automatically make a database mistake reversible, a tangled system understandable, or a fork as trusted as the original. That boundary is where the argument about agentic development gets useful.

Convex Abstract — James Cowling with Jean-Denis Grèze, David Cramer, and Theo Browne
Nobody Agrees About Agentic Development · 49-minute panel · Published October 8, 2026 · Source ID: tU0leMxkqzE

Theo Browne says T3 Code would not exist in its current form without agents. David Cramer says Sentry gets internal productivity gains, but the broad reality of production software has not changed very much. Jean-Denis Grèze says a category of work that helped make Plaid defensible has become routine at his new company, Town. These are not three estimates of the same productivity multiplier. They are three different accounts of where the value moved.

The disagreement becomes sharpest around architecture. Browne sees a new freedom to build, measure, discard, and learn. Cramer sees missing context and duct tape accumulating faster than people can understand it. Grèze thinks experienced engineers need to recalibrate their definition of good architecture, without pretending that every mess can now be repaired cheaply.

A useful reading of the panel is that cheap implementation changes the cost of some errors, not the cost of all errors. The distinction reaches beyond code review into junior development, incident ownership, open source, and the question of what remains defensible when a competitor can produce a working copy.

These views come from builders with different businesses: Cramer co-founded Sentry; Grèze, formerly Plaid’s CTO, leads Town; Browne builds T3 Code, T3 Chat, and T3 Stack. Moderator James Cowling is Convex’s co-founder and CTO. Convex hosts the discussion, and Browne is introduced as an early supporter of the company. Their product examples are operating experience and commercial perspective, not a controlled comparison of agents.

The gain is not necessarily more code 02:38 ↗

Cramer separates internal efficiency from product transformation. Sentry spends a substantial amount on frontier models, he says, and better access to information helps people become more self-sufficient. But more people shipping code does not automatically mean more software with users, more meaningful production workloads, or a fundamentally different demand for error monitoring. His two promising directions are doing R&D that previously was not possible and reaching an expanded market of builders. He does not claim that the existing production world has already been remade.

Grèze cannot offer a clean before-and-after productivity measurement either: Town was built from the outset in the newer model environment, after a pivot he dates to the previous November. Instead, he compares kinds of work. At Plaid, bank integrations required strong engineering and were difficult and expensive. At Town, he says, integration contracts can be formalized, and agents do that implementation work. The point is not that every bank integration or every Plaid advantage has vanished. It is that work which once carried scarce engineering value need not be the moat of this new business.

His second gain comes from removing handoffs. Product development used to require an engineer, designer, and product manager to communicate and align. Someone with all three sensibilities can now iterate on a feature without waiting for those exchanges. Town also relies on existing infrastructure such as Convex and Sentry rather than rebuilding systems underneath the product.

That makes discovery faster, not necessarily sturdier. Grèze describes getting to product-market fit for a feature in a more brittle way, then stepping back to do proper engineering. Agents can undermine abstractions, add unnecessary layers, or introduce another framework. The panel’s joke about an agent changing the tests to make them pass has a serious implication: the system doing the work should not have unrestricted power to redefine the evidence of success.

Browne’s answer is more categorical. Maintaining a large open-source project with a two-person team and roughly a thousand pull requests would not be viable without agent triage. Users can contribute fixes, propose features, and fork their own variants; agents help maintainers find contributions worth examining. The new capability is an operating model, not simply faster typing.

Review the risk, not an ideology 08:01 ↗

At Town, Grèze draws a line around safety-sensitive changes: API boundaries, authentication, storage, encryption, keys, and access controls. Those trigger blocking review. For product-layer work that leaves those foundations alone, evaluating behavior can matter more than reading every implementation detail.

He gives a concrete example: Devin provides short recordings of a feature working, letting someone assess the product without repeatedly clicking through it. Strong CI and automated tests catch other classes of mistakes. This is still testing. It is a change in how evidence is presented and where scarce human attention goes, not a claim that a feature recording proves security.

Cramer agrees that consequence and reversibility matter. A significant database change is not something to merge unread. An agent cannot reliably act on constraints that remain in its operator’s head, and specifying every nuance can consume more time than the tool returns. His practical calculation is whether a mistake matters and how quickly it can be unwound. Sentry still requires code review for production systems, even though he would like to relax the rule in some cases.

Cowling adds a different defense of review. Convex reviews all its code, and correctness ultimately remains the author’s responsibility. Review is not merely a bug-catching ritual: it evaluates architectural choices and develops engineers. Abolishing it requires an answer to what will replace those learning functions.

Browne’s challenge is to define an acceptance bar rather than move it every time models improve. At what point would an engineer stop reading the code and accept the agent’s recommendation to merge? That question does not settle where the bar belongs. It exposes the difference between an explicit standard and a refusal that can never be satisfied.

A thousand PRs need a different safety system 11:39 ↗

Browne disputes the idea that reading every diff is a reliable way to find the hardest bugs. In his experience, he has approved more reviews containing bugs than caught bugs simply by inspection. His proposed response to a tenfold increase in output is not tenfold review time. It is making failures cheaper: detect them earlier, fix them faster, and limit the consequences.

He reports that T3 Code receives around 1,000 PRs every two weeks, and that the team merged more than 100 in the preceding 24 hours without major regressions. Those are his reported results, not a general guarantee. He attributes them to gradually built systems: early detection, agent testing and review, and an architecture that gives each change a clear, logical place. Without those boundaries, agents plug gaps in strange ways.

For understanding a growing codebase, he recommends asking agents to explain it. The benefit is an effectively inexhaustible supply of questions without repeatedly interrupting the senior engineer who built the system. But an explanation can be wrong. His prescription includes back-and-forth questioning and second opinions, not treating one fluent answer as established truth.

The critical distinction is between delegating a task and removing the system that makes delegation safe. Browne’s throughput story depends on that system. It is not evidence that every team can skip review because models write plausible code.

Cheap experiments, expensive spaghetti 13:50 ↗

When Cowling asks how architectural judgment develops, Browne’s answer is to get it wrong repeatedly and gradually learn what is right. He thinks that has become easier. Cramer explicitly disagrees.

Cramer’s objection is about composition. Agents may design an isolated component reasonably well while missing the constraints that make it fit the whole system: scale, user counts, which requirements exist today, and which are future possibilities. Software design also depends on maintaining context across repeated rebuilding. Faster implementation can compound glue and duct tape, while the scope grows beyond what anyone understands. The problem is not just a bad function; it is a system whose moving parts interact badly.

Browne counters with his own part-time, solo cloud work. He has been generating load tests to compare several architectural ideas, including one proposed by a model. An experiment that once demanded perhaps 5,000 lines of code and months of effort could now be built, measured, and discarded. Previously, discovering that it was slower than expected would have been demoralizing enough to discourage the attempt. Now the cost of being wrong about an experimental design feels close to zero.

That example demonstrates why the disagreement should not be flattened. Browne is describing an expanded ability to test hypotheses. Cramer is describing the cost of accumulated decisions in an integrated system. Low-cost exploratory code and low-cost production correction are not the same proposition.

Grèze challenges the older standard underneath both arguments. In the previous model, giving a junior engineer a difficult two-month assignment could produce something that failed to scale and was expensive to change. A senior engineer would help with design because preventing the wrong architecture was cheaper than replacing it.

Agents let a junior take on more complexity, but also produce more architectural drift. Some of that drift is now cheap to correct, so experienced engineers may react too early to code that offends old sensibilities. Yet there is still a threshold where complexity becomes so tangled that even an LLM struggles to repair it. Grèze does not erase that threshold; he says the judgment about where it sits has changed.

Grèze introduces a question he says he does not yet have an answer to, then proposes a model-centered definition: good architecture lets an LLM meet product requirements with less human intervention, rather than being judged primarily for humans. That is not permission to optimize only for the next passing test. He asks about the next ten features, whether future agents can follow the company’s way of building, and whether tests invite a model to take shortcuts such as returning true everywhere. The discipline survives even when its answers change.

Curiosity is leverage; ownership is still work 17:32 ↗

Cowling worries that the industry is consuming judgment it has not learned to replenish. He points to what he sees as the frontier labs’ preference for already-developed senior engineers and compares it to remote work during COVID: executing existing plans could look highly productive, while making the next plans proved harder. His hiring characterization is an observation in the discussion, not evidence of a universal company policy. The underlying question is whether teams are benefiting from accumulated architectural wisdom while weakening the process that creates it.

Browne sees extraordinary access for a particular kind of young developer. Some students in his community were already ahead of their teachers and isolated from stronger peers. Stack Overflow and an old technical talk with a bad microphone were imperfect substitutes for someone who could answer their questions. Agents let them build toy systems, explore large codebases, and test ideas directly.

He says more than 85% of his audience is at least 25, but two of T3 Code’s top five contributors are 16. He explicitly distinguishes this intensely curious minority from developers generally. It is not evidence that every junior automatically becomes a senior.

His concrete learning exercise is better than a vague instruction to “use AI”: ask what the maintainer repeatedly rejects, examine merged and rejected changes, and learn how to avoid the same mistakes. The tool can reveal the local standards of a project to someone who does not yet have a professional network.

Cramer is less persuaded that the human pattern is new. He began contributing to open source at 15, before LLMs. His more important example comes from recent Sentry incidents: poor postmortems were not caused by a lack of tools, but by missing ownership or technical capability.

A weak postmortem says to fix the bug. A stronger one asks how to redesign the system or add verification layers so that the same class of outage cannot recur. That difference requires expertise and responsibility. An agent can help implement either response, including the inadequate one. Generating a polished explanation does not supply the ownership that the incident demanded.

Discard the artifact without discarding the thought 27:31 ↗

Cowling distrusts AI-written design documents partly because of his own psychological anchoring. Once a document or some code exists, he starts treating it as reality. At Convex, he wants people to struggle with the problem first and develop a sense of its difficulty before a generated artifact supplies an apparently finished answer.

Browne experiences the opposite effect. Cheap generation makes throwing work away liberating. He no longer feels pressured into “guilt-merging” a nearly right contribution after its author has spent days revising it. If agents produced the implementation, rejecting it feels less costly.

He also describes a way to retain useful effort without accepting the whole artifact: have an agent find related open PRs, cherry-pick useful commits, and credit the contributors in the new PR description before closing the old ones. The idea is disposable implementation, not anonymous appropriation.

His volume is intentionally excessive: he estimates producing around 150 documents and HTML pages daily, reading roughly half, and losing track of 95% or more. After two regressions, he may request three architectural alternatives, skim the proposals, choose one, and ask for a PR with a way to inspect the result. For every PR he lands, he estimates abandoning five or more.

These are different psychological risks, not a contradiction that must be resolved in favor of one personality. Cowling fears mistaking an available answer for a considered one. Browne uses a surplus of available answers to reduce attachment. The meaningful test is what understanding or observable result survives after the draft is thrown away.

You can copy the code. Can you copy the customer relationship? 30:32 ↗

Asked what prevents replacement by a competitor with more tokens, Grèze rejects certainty: nobody knows where the enduring moats will settle. Possibilities include the customer relationship at the application layer, hard infrastructure problems, and network effects between agents. A computer can operate a competing product where it performs well, add other capabilities around it, and itself become something another computer operates. His token-arbitrage thought experiment challenges the idea that a product boundary automatically remains a business boundary.

He also questions the economics. Traditional software could support margins around 70–80%; an LLM product reselling tokens with a markup need not have that profile. These are his broad comparisons, not audited figures for the panelists’ businesses.

For Town, he favors owning the user relationship. His prediction is that people will not want 17 separate AI products: they may use one or two top-level interfaces that operate coding tools or spreadsheets underneath. He likens those interfaces to browsers or devices more than apps. He imagines that developing over two or three years, potentially changing information distribution and advertising economics. It is a forecast, not an established transition.

He also looks to network effects. Moving one person off a tool differs from moving a whole team out of an agent-to-agent framework. At the same time, cheaper data migrations can weaken older forms of SaaS lock-in. His furthest speculation is a world in which the company’s central activity becomes allocating capital to intelligence that turns energy into economic output. The panel does not establish that outcome.

Cowling brings the question back to distinctive talent and domain judgment. Browne answers with distribution: his community brings skilled contributors, users, and attention. A closed-source competitor may not attract the same interest; an open-source competitor exposes ideas T3 Code can also adopt. A fork can reproduce implementation more easily than it can reproduce his reach. He acknowledges a threat from lab tools becoming good enough that users stop moving between models.

Cramer likewise emphasizes distribution and spending power, but rejects “moat” as an impermeable barrier. He recalls telemetry customers switching vendors in a month because the savings justified it, even before AI. Slack’s network effects are a different source of inertia from merely owning code. A startup without customers does not yet have a moat to defend; it is trying to get started.

One subtler cost of easy copying is lost feedback. Cramer values pull requests even when he dislikes them because they reveal product demand. If a user silently forks and solves the problem alone, the maintainer may never hear about it. Customization becomes easier while the shared learning loop can weaken.

Open source gets the risk first—and some benefits first 38:59 ↗

Cramer does not explain public development solely as a competitive tactic. He likes the community and the act of building in the open. Copying was always a risk. He also makes an important licensing distinction: Sentry uses a fair-source license, and he says legal restrictions protect its IP. Source visibility does not mean unrestricted permission to run a competing business. He says Sentry has not yet needed the litigation he threatens.

Browne distinguishes exposing a backend, server, and system implementation from publishing a client SDK or package. Full backend source makes it easier for a capable agent to search for exploitable defects and reproduce behavior. His argument for staying open is not that those dangers are imaginary.

Instead, he predicts that closed systems will also become vulnerable to automated discovery and cloning from their externally visible behavior or API definitions. In that account, the uniquely open-source disadvantage is temporary because the wider problem gets worse—not because security gets solved. Open projects face it sooner and may develop defenses sooner.

He reports a compensating effect at T3 Code: thousands of people inspecting the code and regularly contributing fixes for issues that could be exploitable in specific circumstances. That is his account of community security work, not proof that every vulnerability has been found.

The product benefits are concrete. Agents can inspect, understand, modify, and customize accessible source. Users can adapt the software, report bugs, and gain confidence from seeing how locally running code works. Browne expects those advantages to outlast the period of asymmetric exposure.

Cowling says Convex customers pay for trust even when others clone its source. But he also points to frontier labs’ secrecy around valuable models and training infrastructure as a counterexample to the idea that openness always wins. The panel leaves the tradeoff unresolved: transparency can improve adoption and inspection while simultaneously lowering the cost of attack and imitation.

Three career prescriptions, not one reassurance 44:36 ↗

Browne tells young engineers to follow curiosity. Agents do not provide the curiosity; they let a person act on it. Rust projects and his Lakebed cloud work moved from things he might someday have time to attempt to experiments he could start now. He sees designers becoming developers, developers exploring design, and marketers building software because the old implementation barriers no longer halt exploration so early.

Grèze focuses on the outcome for another person. Writing code is not what he misses or what gives the work its value. Delivering something that makes someone’s life better is the satisfying part, and that opportunity is easier to pursue.

Cramer offers less fashionable advice: do not start a company immediately; work with people who know what they are doing. Browne agrees that startup-building teaches a particular set of skills, not necessarily the craft of making excellent products. Cramer describes hiring people with unimpressive résumés because their public projects showed care. His own pattern is building publicly, taking feedback, caring about results, and working alongside skilled people.

For builders, the practical lessons are specific:

Those lessons are an editorial synthesis of a disagreement, not a framework all four participants endorse. Cowling’s closing definition provides the shared ground: engineering means solving problems for people in the presence of constraints. Agents change those constraints. They do not make every constraint, or every consequence, disappear.

Source and method. Written from a full automated English transcription of the original English audio; non-English auto-caption text was used as a secondary cross-check. Speaker names and roles follow the publisher’s metadata. Numerical results, product practices, and predictions are attributed to their speakers; they are not independently measured benchmarks. Timestamp links return to the source. No reader-specific personalization has been added.