The trust ladder for agents
Definition: a way of describing how much independent work you hand a coding agent, where the amount you can hand off grows only after the agent proves it can check its own output and the codebase makes a wrong move hard to take. lauren-tan uses the ladder in 2026-10-05-youtube-poteto-software-factory to explain how she went from babysitting one agent to letting many merge without her. The host opened the stream with it, “increasing your velocity and climbing the trust ladder with agents so that you can ship more and more and more” (0:26).
What holds you low
Her diagnosis of what caps the rungs is not model capability. Once a frontier model is good, “the bottleneck is no longer the agent,” it becomes your ability to express intent so the agent can carry it out (7:17). See matt-pocock’s parallel line about finding the exact word the agent then reinforces in its own traces. Two things hold you low:
- Frontier models still take the easy shortcut (6:04): “even the frontier ones tend to take shortcuts. They tend to do the easy thing.” So the way up is to make the easy thing the right thing, by changing the environment. See environment-and-constraints.
- Low trust forces micromanagement (26:31): when you cannot trust the work, you fall back to locking in and supervising every step, and that mode consumes the time you would spend building the skills and tools that would have raised the trust. The first rung she ever cleared was agent-verification-skill, because it took her out of the proxy role between the agent and the result (17:03).
The top rung
At the far end, agents merge their own pull requests overnight while she sleeps, and she reviews the commit history in the morning and course-corrects (52:02). Getting there “is very hard,” and she refuses to sell it as a button (51:20). It is the payoff described in michelin-kitchen-metaphor as the open chain restaurant: build one good kitchen, then be able to walk away from it.
A ladder, so trust is reversible
The word ladder matters. Trust is built rung by rung through evidence, and it can drop when several agents repeat a mistake, which is the trigger to fix the environment rather than the agent. That is the opposite of granting trust up front. Read against one-prompt-awe and ai-psychosis, the ladder is the corrective: those describe over-trust granted from unverified success, and this one puts verification first. She works the same point into her own guide, “pstack does not ask you to trust an agent on day one.”
Related
- agent-verification-skill: the lever that clears each rung
- environment-and-constraints: the work that raises the ceiling
- michelin-kitchen-metaphor: the scaling story the ladder is part of
- one-prompt-awe and ai-psychosis: the failure mode the ladder argues against