I watched Zach Lloyd’s talk this week and recognised my own year in it. I had built the back half of the loop he describes, without the map.

The Thesis

Lloyd, Warp’s founder, gave a talk at the AI Engineer World’s Fair in July, now on YouTube: software engineering is becoming factory engineering. Ideas enter a loop. Agents triage them, write a product spec and a tech spec, implement, review and verify. Review agents or people step in at fixed points. Monitoring feeds the top of the loop again.

You won’t be building the product. You’ll be building the thing that builds the product.

— Zach Lloyd, Warp, slide: The Factory Mindset

Warp runs on it. When it open-sourced its client in April, the usual open-source pain arrived: noisy issues, sloppy PRs, review hell. Warp’s answer was a factory to absorb it, with a public view at build.warp.dev: every issue in flight, and which agents and contributors are working on it. Lloyd predicts every sizable project will run one, “kind of like the way that CI/CD became just like, oh, of course you have that.”

His architecture slide treats the coding agent as a part. Claude Code and Codex sit side by side as harnesses, the programs that run a model against your code. Warp builds the orchestration around them.

The Playbook: Backtest Everything

Since June, every change I make to a critical path gets replayed against real history before it ships. A prompt edit, a model swap, a matching rule, a ranking tweak. Agents write the change in minutes. The backtest decides whether it ships. Across FameCake, Pricogni and other systems, the same techniques keep paying off:

  • Replay real history, read-only. A FameCake moderation prompt change ran over 695 real posts from the last 90 days and was scored against the verdicts already stored. It caught 10 of the 11 posts a person had rejected.
  • Run old and new on the same inputs. Two worktrees, two builds, or several commits in one pass. One run took 23 million records through main and two candidate commits. The candidates agreed on every entity, and 826 entities changed.
  • Hide a label you already have. Pricogni matches products across retailers. 60,026 listings carry a barcode, which is the right answer for free. Hide the barcode, run the matcher, score it. Precision went from 77.6% to 90.2%.
  • Run the same prompt twice first. Models are not deterministic. A control run of the old prompt shows how much verdicts move by chance before you trust a difference.
  • Read the changed rows. One rule change showed 67 more matches and looked like a clear win. Reading the rows found a record the rule would replace with the wrong one, so it would silently disappear.
  • Price the result in decisions. For a FameCake alert on offline screens, the backtest priced each threshold in false alarms per month: 595 at 16 hours, 147 at 24 hours. The choice took one minute.
  • Kill what loses. A classifier prompt change moved 2 of 137 uploads, and one of the two moves was wrong. It did not ship. A model that scored 88% against 87% at 8x the cost did not ship either.
  • Get a second opinion on the backtest. Once, a backtest reproduced a bug because its simulator copied the flawed logic. A review by a second model caught it. A backtest that copies the code copies its bugs.

The best backtests stop being one-off scripts. Pricogni’s matcher backtest now runs weekly as an audit: 400 sampled items, barcodes hidden, no writes. The health check fails if precision drops below 85% or the audit is more than 14 days old. It replaced the monthly human sample. That is the loop closing.

The rule under all of it

Expected results come from the running system or labels it already holds, never from the agent. An agent that writes both the change and the expected output has graded its own exam. StrongDM gets there differently: it often keeps its test scenarios outside the codebase, where the coding agent never sees them.

From Scripts to a Factory

The same rule is now moving into CI across a few dozen repos, where agents write most of the changes in parallel. Nobody planned a factory. Each piece came from something that went wrong:

  • One config for every agent. Shared skills in every repo, plus a hook that stops before a force-push, hard reset or broad delete. The instructions say a check that never ran is not a check that passed.
  • Releases that ship what was tested. Each service builds once. Promotion moves the exact commit that staging tested, with no merge in between. After a deploy, every route must report that commit. A superseded run once finished green having deployed nothing, so production now requires a staging run that actually deployed.
  • A board that moves itself. Every PR links a card. Deploys move cards forward, never backward. Shipping to staging posts a preview of every change first.
  • Golden replay. The backtest playbook wired into the release path. The harness runs the current production build beside the candidate and sends both the same real requests, anonymised. It calls production twice, so fields that change on their own count as noise. Any other difference fails the release.

Golden replay is the newest piece and still in draft. Its first dry run pointed both sides at the same build to test the noise detection: 20 of 20 requests matched, and the only noisy field was a trace ID. The backtests above already caught real bugs, though. One sync change, run on the old and new builds against a broken input file, showed the old build deleting 2,999 of 3,000 rows. The new build deleted none.

Same Loop, Opposite Ends

Lloyd’s loop has verification in it. He started at the front, and that makes sense for an open-source project buried in issues. I started at the back, because agents already wrote more code than I could prove safe. Karpathy put it in one line last November:

Software 1.0 easily automates what you can specify. Software 2.0 easily automates what you can verify.

— Andrej Karpathy

His slide makes my case for me. If the harness is a swappable part, then the code it writes is not where the factory’s value sits. I can swap the coding agent. I cannot swap the checks. The factory’s product is not code. It is proof.

We agree on self-improvement. Lloyd wants observer agents that watch a skill fail and rewrite it. Mine are cruder so far: lessons from each backtest go into the agent instructions and memory, so the next change starts from them. One threshold decision was saved “so nobody re-litigates.”

What’s Still Missing

  • The front half is manual. Triage and spec agents are still a plan. build.warp.dev is ahead of me there, though Lloyd is candid: “It’s not working perfectly, but it is working.”
  • People approve every merge. StrongDM’s rule is “Code must not be reviewed by humans.” I am not ready to bet production on that. The backtest makes the approval fast, not optional.
  • Replay proves agreement, not correctness. A backtest shows the new version behaves like the old one, or like a label. If the old one had a bug, the new one passes with it. An intended change needs a reviewed new expectation.
  • Throughput is a trap. Lloyd wants factories measured by how much software they ship, and what it costs in human time and tokens. Measure only that, and a weak gate becomes a fast way to ship bugs.

Lloyd’s slide asks whether this is a bummer, and answers: “You’ll code less, but you’ll ship more.” I think that is right. But the shipping part was always CI. The future of AI is CI.

P.S. If you want to feel the back half of the factory, play MVP Mayhem. You are the gate: agents ask, you approve or deny. In testing, a bot that read each command won in under seven minutes. A bot that approved everything was dead in about one.