Verifiability and one-way doors
Definition: the ceiling on how much you can let agents merge on their own is set by whether the work can be verified. Where checking is possible, a mistake stays cheap to undo. It is the hard limit on 2026-10-05-youtube-poteto-software-factory, and the host raises it deliberately.
Two-way doors and one-way doors
The host’s framing (56:04): “Some PRs are like two-way doors, you can merge it and then revert it, cheap to walk back. Some PRs are one-way doors, that will cause data loss of some kind, or do something that can’t be easily walked back.” He asks what to do when most PRs are one-way doors, naming medical apps, law, and finance as security-conscious, hard-to-reverse settings.
The answer is verification
lauren-tan’s reply puts it back on the same lever as everything else in the stream (56:32): “It all comes back to the quality of the verification that you’re able to get out of your agent. For domains where the work is verifiable, this is easier, and the one-way doors become two-way doors. If you’re working on something that is very hard to verify programmatically, then you’re in a position where it’s very hard to get to that point.” So verifiability of the domain is a precondition for the trust-ladder-agents, not a side note. See agent-verification-skill.
Where it holds, and the formal case
She calls software quite verifiable in many cases, and mathematics partly so, with a written proof as the example (57:24). Her prediction is more agent-oriented languages that make verification cheap, and she names bend, a language that “marries programming with proofs,” where you formally verify that code is correct and the compiler rejects any change that breaks a declared law (57:56). The old way was to write the proof in a separate language, Lean or TLA+, and hand it to a solver to confirm the cases were checked and to catch a race condition. Her summary: if it compiles, and the proofs show it is correct, why would you not merge it (59:03).
What she will not oversell
She answers the door question by admitting she does not have it solved (57:38): “That’s a great question that I don’t really have the answer to, and it’s something the industry and us as engineers will have to figure out. Not all domains are verifiable.” Earlier she had already refused to sell the destination as easy (51:20): “It’s very hard to get to this point. I don’t want to sell this as something you can just do easily by using pstack. It takes a lot of time to look at where your agents fail and thoughtfully set up guardrails and constraints so they do the right thing by default.”
Read against ai-psychosis and one-prompt-awe, this is the boundary of the whole method: environment and verification are what let you step away, and only up to what the domain can actually prove about itself.
Related
- agent-verification-skill: the mechanism that turns one-way into two-way
- bend: proof-carrying code as the strongest form
- constraint-driven-codebase: constraints make a whole class of breakage unrepresentable
- trust-ladder-agents: the ladder has a ceiling set by verifiability