Agent verification

Definition: giving a coding agent the means to run its own output and observe the result, the “hands and eyes” that let a loop close without a human standing between the agent and what it produced. lauren-tan ranks it above everything else in her toolkit in 2026-10-05-youtube-poteto-software-factory:

[16:24] “Even if you don’t use pstack or my skills, I think that the single most important skill that should be in your toolkit is verification.”

What it is

The analogy she gives (16:35): “give your agent hands and eyes. The agent is able to run the code, and actually interact with it like a normal human user would, and also do things like debug it, take traces and snapshots.” In pstack the concrete artifact is a small CLI over Playwright and the Chrome DevTools Protocol, and the agent uses it to drive the real app. See determinism-extraction.

The part that makes a loop a loop

A claim about the word “loop” (17:50): “the most important part of a loop that allows it to be a loop is the verification part, because the agent is able to verify its own work, and that takes you out of the equation.” Her other skills, the how skill and the unslop skill, were good prose but left her as the one who looked at the output and noticed when something was wrong. An agent that cannot observe the result of its work has no path back to iterate. Verification was the first skill she built at cursor and the one that let her climb the trust-ladder-agents.

Autopilot and fuzzing

Turn on full autopilot and it “triggers off this very intense, rigorous verification loop where it will spawn a bunch of verifier agents for every pull request and it will fuzz” (52:50). Fuzzing here means running the application and clicking around like a real user, hunting for regressions and bugs in the implementation, then fixing them and repeating until the pull request can land. It is token-intensive and tunable, from ten verifier agents down to one. It also became shared infrastructure: every app at Cursor or xAI carries an auto-maintained verification skill, and the meta-skill /create-verification-skill walks an agent through writing one for a new app (19:10).

Verification is how one-prompt-awe and ai-psychosis get answered operationally. The ai-weekend pages treat unverified trust as the danger; this is the machinery that puts verification first, so trust rests on a checked result. See verifiability-and-one-way-doors for the limit, since a domain that cannot be verified programmatically caps what verification can deliver.