The Factory Has to Improve, Not Just the Agents

Suraj Gupta, Warp — AI Engineer — “How Software Factories Improve Themselves,” TN3mj92oZ8I


Self-improving agents and software factories have been said separately all week. Suraj Gupta, who leads harness development at Warp, wanted the missing combination: how the factory itself gets cheaper, faster, and less wrong over time. He is not a bystander. Warp started as an agentic development environment born from a terminal, claims about a million active users, and is now selling a cloud agent platform — he called it Oz on stage — for teams building factories. Treat the demos as product. Treat the three loops as the argument.

The job, as he framed it, is shifting from building products to building and maintaining factories: automations that take work from triage to production. Humans are supposed to phase out of the improvement loop so it can run automatically. Skills go stale. Agents rediscover root causes they already paid tokens to find. Opus prices get applied to “the issue is a duplicate.” Those are the three leaks he plugged, in order: skills, persistent memory, model routing.

An outer loop that opens a PR

Skills are procedural memory: a procedure you expect the agent to run over and over. A triage agent gets a skill for reproducing issues. Then humans leave feedback, the agent’s own trajectories teach it things, and the markdown you wrote last month is dated.

Warp’s fix, which Gupta said has been productive in their internal factory, is a second agent. The inner loop applies the skill — it does the triage. The outer loop watches those runs, looks for mistakes and human feedback, and updates the skill. Their client repository is open-sourced and run as a factory. A GitHub issue comes in. The triage skill’s job is to notice what the reporter did not provide, whether this should be built, or whether it duplicates an existing issue. A simple workflow fires on every new issue. The outer-loop agent then reads the inner agent’s decisions and collects signals: thumbs up or down, user comments, Warp employees commenting on what it did. Synthesis becomes an update to the inner-loop skill.

Step four of the demo: add the triage skill and open a pull request. Improvements are tracked in Git. A human reviews the skill change so the outer loop cannot silently make triage worse. That is the whole point of putting self-improvement on a PR instead of letting the agent rewrite its own prompt in place.

A fact store the next Sentry run can actually use

Skills remember procedures. They do not remember that last month’s Sentry incident was the same root cause. A Sentry agent gathers context, fixes it, and the next similar issue may get lucky — or it may burn the same tokens rediscovering the same cause. Persistent memory, in Gupta’s telling, is a fact store scoped to an agent. Same outer-loop pattern: extract facts, learnings, and outcomes from the inner run so later runs can lean on them.

On Oz, every cloud agent can attach a memory store. He showed a Sentry agent that had collected memories about past root-cause work. Humans can version the store, add memories, delete them. Creation is primarily agent-driven. You can see where a memory was sourced — he tried to open a run live; staging was IP-gated and his address kept changing, so the click failed on stage — so you can drop a local maximum you do not want in the store. On a later triage run, the same agent reported which memories it used. In the example, five of them sped up the work.

In Oz, he said, this works across harnesses: Warp’s own, Claude Code, Codex. Memories get created regardless. You still review, version, edit, delete. Fully traceable. That portability claim is Warp’s product claim, not an independent benchmark.

Stop paying Opus to file duplicates

Routing is the third improvement loop because a factory that always calls Opus for triage and simple CI fixes becomes prohibitively expensive. Organizations have already been burned. Warp ships “auto models”: out-of-the-box routers they keep re-evaluating as new models land, aiming at Pareto efficiency, so a new user is not pinning everything to Opus or Haiku.

You can also write your own rules. Gupta dragged over a config: database migrations on GLM, runbooks and API documentation on Qwen. Classes of tasks mapped to models you have actually seen work. At this point, he admitted, that mapping is more art than science. Next they want customer-facing evals — knobs on those configs, not a generic public benchmark — so you can see whether Haiku, GLM, or Opus wins on your class of issues.

Internally Warp already runs an eval sidecar beside the routing rules. Best-of-k: a prompt fans out agents in Oz across models. Their finding, as he stated it: UI tasks run well on GLM; you do not need Opus. More efficient. They intend to put that into the product so customers can do the same in their workflows.

That was the talk. Booth on that side of the hall. Come talk. The concrete takeaway is not “self-improving factory” as a slogan. It is three loops you can actually implement: an outer agent that PRs skill updates a human can reject, a versioned fact store so Sentry does not start from zero, and a router so triage is not an Opus bill.