YouTube: Poteto on shipping thousands of PRs with agents

Source: YouTube live stream MN9dGgmLyso, “LIVE: Poteto (creator of pstack) on shipping 1,000’s of PR’s a month at SpaceX” https://www.youtube.com/watch?v=MN9dGgmLyso, streamed 2026-10-02 on the Matt Pocock channel https://www.youtube.com/@mattpocockuk. oEmbed record read 2026-10-05: title as above, author “Matt Pocock”. Host matt-pocock talks with lauren-tan (X handle @poteto), creator of pstack. The description reads “Matt Pocock talks with Lauren Tan (poteto) about how to ship extraordinary amounts of high-quality work using agents.”

Ingested: 2026-10-05, from the auto-generated English captions of the full stream (about 1 hour 5 minutes). The captions garble proper nouns. Each correction below was checked against a primary I read before filing, and the read is recorded in the linked pages’ Sources section:

  • The guest is Lauren Tan, X handle @poteto. The captions render the handle as “potato” and once spell a GitHub path as “potato-noodle n-o-d-l-e”. Her GitHub profile https://github.com/poteto and X https://x.com/poteto confirm poteto, not potato. The oEmbed title and the stream description both say poteto and pstack.
  • The skill library the captions call “PAC” is pstack, upstream at cursor/plugins/pstack https://github.com/cursor/plugins/tree/main/pstack. See pstack.
  • The product garbled as “grockbot”, “graphbot”, “grabbot” is Grok Bot, xAI’s product. Her LinkedIn bio reads “Building Grok Bot and Cursor at SpaceXAI”; her pstack guide Pt. 1 (https://x.com/poteto/article/2094457600259842065, Aug 31) says she “started working on Grok @Bot about 2 months ago”; the xAI Grok Bot Marketplace (https://x.ai/bot/marketplace) lists “dr eggbot by Lauren Tan”. See lauren-tan.
  • “Uncle Bob” (0:05) refers to the prior stream on the same channel, “LIVE: Uncle Bob on Software Fundamentals in the Age of AI”. No page filed.
  • “the labs talk about” hill climbing (18:17) and “Andre Carpathy” releasing “auto research” (18:35) resolve to andrej-karpathy and his autoresearch repo https://github.com/karpathy/autoresearch, published March 2026. See that page for the train.py, program.md, five-minute fixed budget, and “super lightweight skill” details.

Quote provenance: quotes come from the auto-generated captions. False starts, [snorts], [laughter] and cleared throats are trimmed. Timestamps are the caption clock in minutes and seconds, ascending. Speaker turns alternate between host and guest; the captions use >> for the guest’s turns and do not label the host, so attribution below follows the turn markers and the question-and-answer shape of the exchange.

Summary

Lauren Tan describes how she got to roughly 2,500 pull requests merged to production in a month at SpaceX AI, and argues the number is a by-product of two things people skip: building the environment the agent works in, and giving the agent verification so it can check its own work. She reframes the buzzword “software factory” as the Michelin kitchen, because factory does not carry craft. The trust ladder is how she describes going from micromanaging one agent to letting many run unattended. She runs two loops, an inner loop that writes code toward a snapshot of intent and an outer loop of connectors that pulls context in from Slack, X, and Linear. On skills she agrees with the host that a skill is process turned into words, skills as process, and expects them to shrink as models improve.

The trust ladder and the intent bottleneck

Tan’s route here starts from micromanaging a single agent on a burnout side project, which became the seed of pstack. She joined Cursor in March, worked on the laggy agents window, and became the “meat proxy” between the agent and Chrome DevTools (5:09 to 5:13). That annoyance drove the whole method.

  1. The bottleneck moved off the model (7:17): “as the models get really really really good it almost becomes like the bottleneck is no longer the agent. It becomes your ability to express your intent and your goals in a clear way that the agent can understand and actually carry out.” See trust-ladder-agents.
  2. Domain expertise is more useful now (6:51): “I actually feel like domain expertise is more important than ever.” Her case is that a doctor or lawyer who is slightly tech-curious and holds a clear vision can build the product, because the transfer of intent is now the hard part.
  3. Frontier models still cut corners (6:04): “even the frontier ones tend to take shortcuts. They tend to do the easy thing. A lot of the skills that I’ve built have been around: how do I make the easy thing the right thing?”
  4. Language is the lever (8:42 to 10:04): the guest describes what she loves, “when you find a word that the agent hooks on to and goes, okay, I’m going to reinforce that word, I’m going to reuse that in my thinking traces,” with TDD as the early example and the host’s grilling as a case where it works. The host’s own tip on eliminating tautological tests, which she says she copied, is what she pauses on: “there’s a lot of meaning to that word. It’s almost like compressed. You compress a lot of intent and meaning into words.” Both rest on the same bet, skills-as-process.

Verification as the core skill

The host turns to what people can change today. Tan names one thing.

[16:24] “Even if you don’t use pstack or my skills, I think that the single most important skill that should be in your toolkit is verification.” See agent-verification-skill.

  • What it is (16:35): “give your agent hands and eyes. The agent is able to run the code, and actually interact with it like a normal human user would, and also do things like debug it, take traces and snapshots.”
  • Why it is the loop (17:50): “the most important part of a loop that allows it to be a loop is the verification part, because the agent is able to verify its own work, and that takes you out of the equation.” Other skills, the how skill and the unslop skill, left her as the proxy between agent and output. Verification was the first skill she built at Cursor and the one that moved her up the ladder.
  • Hill climbing (18:17): with a rubric and a loop, “you can have an agent continually try to make improvements.” She credits the labs with the term and points at Karpathy’s autoresearch as a nearby release. See andrej-karpathy.
  • It became team infrastructure (19:10): “every app that Cursor or xAI has has a verification skill that is auto-maintained. It has become critical infrastructure for our team because everybody uses it.”
  • Autopilot (52:50): turning on full autopilot “triggers off this very intense, rigorous verification loop where it will spawn a bunch of verifier agents for every pull request and it will fuzz. Fuzzing meaning that it will actually run the application, click around, and try to use it like a real human, look for regressions, look for bugs in your implementation, fix it itself, and eventually get the PR to a state where it can land.” Token-intensive, and tunable from ten verifier agents down to one.

The determinism gradient, and tools over prose

Tan built a small CLI inside the verification skill. She is blunt that the CLI is not the interesting part (22:13): “it’s just something that interacts with Playwright and the Chrome DevTools protocol and calls a bunch of APIs. It’s just a bunch of glue.”

  • The gradient (21:08): “agents and skills are like a gradient. You have some parts of the work that are entirely judgment-based, putting together multiple pieces of context and thinking, and then you have the more deterministic parts. Refactoring code from one pattern to another is very mechanical. You don’t need an agent to think about it.” See determinism-extraction.
  • Extract the deterministic part into code (21:45): “I try to extract out the deterministic parts and turn that into code, and just leave only the parts that actually require judgment to the agent. I think of the skill as your wrapper, a wrapper with light instructions around how to use these custom tools.”
  • The cost of not doing it (22:56): without the CLI, “the agent would try to verify its work but it would basically rebuild the world each time, and every agent did it differently. It was not just about context usage but speed.” Putting the CLI in the skill means an agent that uses it inherits the work instead of redoing it.
  • The principle to steal (23:42): “how much of your skills and rules could actually be deterministic. How do I make very efficient use of determinism and non-determinism, and let agents shine at the non-deterministic parts, because that’s what they’re trained to do?” The same logic drives migrations through code mods that crawl the AST and transform mechanically. See determinism-extraction.

The environment is the job

[26:17] “I almost feel like the new job of the engineer is really to spend time on the environment.” See environment-and-constraints.

  • Sharpen the knives (27:25): low trust means micromanaging, and micromanaging consumes the time you would spend improving the setup. “It’s like you haven’t spent the time sharpening your own knives. If you have a dull knife, everything is going to take a long time.” The garlic-masher aside (28:13) makes the same point about tools.
  • A good codebase is a narrow codebase (25:26): “a good codebase is a codebase that’s easy to make changes in. That means you have a lot of guardrails, and the agent or the human is constrained to very narrow paths.”
  • Constraints are type narrowing (29:22): her TypeScript talk background shows up here. “One of my most favorite things about TypeScript is type narrowing. You go from a very broad type that could be anything, and through type guards you narrow the space. There’s a lot of parallels to constraints in your codebase. You’re constraining the number of possible types that can exist.” See constraint-driven-codebase.
  • The internal framework Dune (30:10): a non-open-source framework “kind of like an internal Next.js for our Electron apps”, with “really restrictive lint rules” and conventions such as a directory per feature, so that “it’s actually very hard to write bad code”, and the agent “doesn’t have to think about that anymore.”
  • God files were the trigger (31:29): “the very first couple of versions of the Grok bot were composed of like eight god files which were at least 10,000 lines long, and so I kind of had to break it up.”
  • Turn every observed failure into a constraint (31:46): “This is another important part: observing how agents fail. Every time you see a mistake, every time you see something that could be done better, you step back and think: how do I turn this into a lint rule? How do I make it so the codebase makes this impossible?” The host restates it (32:38) as watching the agent like a hawk, then putting the fix in the environment so the agent “stumbles into the rules” instead of overloading it with things to remember.

The Michelin kitchen, not the software factory

[11:20] “I’ve never really liked the term software factory. Not because it’s not accurate, but a lot of people, when they think factory, don’t equate that with quality or craft, which are very important to me. Michelin Kitchen is the thing I’ve landed on. It’s much more aspirational.” See michelin-kitchen-metaphor.

  • The home cook thought experiment (12:48): as a home cook you are a one-person show. Add your family to the kitchen and “most people would get very stressed.” The question that matters is how to divide work so the sum beats the total of its parts.
  • The chef is the tech lead (14:08): “if you become a chef you’re not cooking all the food yourself anymore. You’re almost like the tech lead for the kitchen, or the CEO of the kitchen: when do you order ingredients, how do you store and prepare them.” This mirrors the engineer who no longer writes the code but owns the outcome, and the reputation still attached to it.
  • Open chain restaurants (34:02): “I’ve spent the time building one kitchen and one restaurant, and now I don’t actually have to be there anymore.” She sees each big project as a restaurant and helicopters between several. The host extends it: manual chats are carrying orders to the chefs yourself; an agent doing expo runs the pass for you.
  • Context not control (41:40): from her years managing engineers at Netflix, “one of the biggest things managers would talk about was context not control. You can drive to an outcome by control, by micromanaging, but what you want is to provide context instead. Teach your engineers to be self-sufficient and then you don’t have to micromanage them.” She says the idea carries straight over to agents.

Two loops, and the triggers that feed them

  • Inner and outer (35:32): “my inner loop is my agent engineers working on the code toward an intent, or a snapshot of my intent. The snapshot can go stale. New information comes to light that I then have to be the proxy for.” The outer loop is everything off the codebase: bug reports in Slack, Linear, or X. See inner-outer-loop-agents.
  • Connect them and you leave the equation (38:37): “How do I take information that my agent needs that I would otherwise have to pass it myself, and teach it how to do it? That removes me from the equation.” She finds the terms company brain and context graph “unnecessarily complex.” The move is to make a trigger, not to build a graph.
  • The concrete wiring (36:51, 42:13): a Slack MCP or a subscription into one channel; the bot watches, and on a bug report it goes and reproduces the issue with the verification skills, confirms the bug still exists on main, and files it into a Cursor project. “Cursor projects are my inner loop and Grok bot is my outer loop.”
  • Coordinator agents (39:10, 42:52): Cursor’s projects feature is “coordinator agents,” one that “has its own computer” and “is a manager of agents. It’s like your executive chef, your chief of staff. It doesn’t do the work itself. It delegates and orchestrates.” On a burst of thirty payloads it picks a topology and distributes the work.
  • Why group, not spawn one agent per bug (44:33): spawning one agent per task “loses the thread between them. You may duplicate work, or you may not think about the higher-level problem.” Multiple slightly different bug reports let the coordinator zoom out and see the real bug is elsewhere.
  • The buffer beats the immediate fix (47:45): an agent scans for React footguns, but she does not tell it to fix each one first. “I tell it to append it to a document, and then every couple of days I look at it and see these are all the same thing. You almost want a buffer, a queue. In pure execution mode you sometimes miss the big picture. A buffer forces you to think about the big picture, and it gives the agent and yourself a way to identify patterns you might miss solving each bug at a time.”
  • Where the PRs actually come from (46:01): “The 2,500 PRs are not 2,500 features. A lot of the work is spent on gardening, another term I really love.” Some of it comes from reading the code, not only from user reports.

Review at scale, and the dark factory

  • Sampling, not tasting every dish (49:34): “You don’t want to be in a position where you’re not tasting your food ever again, but for scale you cannot be tasting every single dish that comes out, especially with multiple restaurants. It becomes about sampling.” Here the factory image fits, since a quality supervisor samples rather than inspecting every item (50:01).
  • Course-correct the environment, not the agent (50:33): “Scrutinize the pull requests rigorously, look at the inefficiencies and bad patterns the agents are doing, and think about how to course-correct the environment, not that single agent. If it was a one-off incident, fine. If multiple agents have the same issue, taking the same shortcut, propagating the same workaround, that’s a sign to amend your kitchen: your skills, your constraints, your lints, your type systems.”
  • It is dark, and it was scary (52:02): “It’s dark in the sense that if my agents are merging their own pull requests it’s become dark. Where I go to sleep, my agents now work 24/7. I have the equivalent of more than ten chiefs of staff, each working on a different area.” She reviews the PR after it lands (52:45): “I look at my commit history, and if I see problems I go and course-correct, revert or modify, add new lint rules.” The first night was “very scary” and took “a lot of bravery.”
  • Not vibe coding (54:49, host): the host distinguishes this from Karpathy’s original vibe coding, “where you forget that code might be a thing,” and reads it as the opposite. “Code and the environment are essential. If they are bad you get bad outputs. Garbage in, garbage out. Maybe there’s a dimmer switch. Some parts dark, some parts light. This is why the kitchen is a better analogy.” She agrees: even a restaurateur “might go take a peek and taste the food.”

One-way doors, and the limit of the method

  • The host’s challenge (56:04): “Some PRs are two-way doors, you can merge it and then revert it, cheap to walk back. Some PRs are one-way doors that cause data loss or do something that can’t be walked back. What if most of your PRs are one-way doors?” She works in a security-conscious setting and names medical, law, and finance.
  • It comes back to verifiability (56:32): “It all comes back to the quality of the verification you can get out of your agent. For domains where the work is verifiable, this is easier, and the one-way doors become two-way doors. If you’re working on something very hard to verify programmatically, you’re in a position where it’s very hard to get to that point.” She calls software and parts of mathematics verifiable, a proof being the example. See verifiability-and-one-way-doors.
  • Proof-carrying languages (57:56): “My hope and prediction is we’ll see more and more agent-oriented programming languages. One of the most fascinating I’ve seen is Bend. That language marries programming with proofs. Proofs is the idea that you can formally verify that some code is correct mathematically, especially if you’ve written it in a functional way. For the longest time you had to write the proof in a separate language like Lean or TLA+, and use a solver to determine you covered all the cases, that you don’t have a race condition.” See bend.
  • Her own admission (57:38): “That’s a great question that I don’t really have the answer to, and it’s something the industry and us as engineers will have to figure out.” She will not sell it as easy (51:20): “It’s very hard to get to this point. I don’t want to sell this as something you can just do easily by using pstack. It takes a lot of time to look at where your agents are failing and thoughtfully set up guardrails so they do the right thing by default.”

Skills as process, and mining your own transcripts

  • A skill is process (1:01:12): agreeing with the host, “a skill is really just process. At the end of the day a skill is just English or language. It’s just markdown.” See skills-as-process.
  • Mine the transcripts (1:00:40, host’s tip, endorsed): “You go and look at your previous transcripts and mine for information. Look through your own prompts to the agents where you correct them, where you have to constantly intervene, and take that higher-level learning and turn it into a reusable skill so that agents stop repeating that mistake.” The host adds that past chats are “a treasure trove of context, because that’s the process materialized. It’s not an abstract idea in your head, it’s the actual thing.”
  • The recall skill (1:03:02): “I have a skill in pstack called recall, which is exactly that.” It came from working on virtualization in the Cursor app, where she kept needing the good context from the last chat in each new one, and collapsing that into a skill meant she stopped writing a long essay every time.
  • Skills will shrink (1:03:47): “As agents get more capable, think of skills as encoding workflows. Skills from last year were about implementation details, like here are the exact script commands you should use. With the latest models you can just delete those parts and focus on the workflow. Over time we’ll see skills get smaller and smaller, more compact.”
  • Your own set of knives (1:01:27): “Everyone should have their own set of knives. Every chef, when they go to a different restaurant, brings their knives with them. Trust is really about trust in your own tools. When you spend the time sharpening them, you can do great things. Somebody might combine Matt’s grill-with-docs or wayfinder skill with some of the execution skills in pstack. Why not read our skills and combine them into your own? Skills are malleable. It’s just language.” The host closes it (1:04:44): “There’s nothing magical in them. If there is any magic, it’s the words chosen and the thinking done to turn abstract process into language. Once that thinking is done, it’s on the surface and you just nick it.”

Touched in the vault

My reaction

The central claim is the inversion: the lever is the environment and the agent’s ability to see its own output, not the model. That is checkable against the vault’s skepticism pages and the two are mostly on different ground. one-prompt-awe and ai-psychosis warn against trusting an unverified result; this stream does not dispute that, it builds the machinery that puts verification before trust. Tan’s honesty about the cost, “very hard to get to this point”, and about the limit, “I don’t have the answer to” the one-way-door case, is what keeps it from being the hype the Jivko video attacks.

Two ideas carry straight into this garden’s own rules. The determinism gradient (put the mechanical part in a script, leave only judgment to prose) is exactly why this vault writes long-lived tools in Rust under tools/ and gates prose on a command rather than a feeling. The buffer beats the immediate fix (47:45) is a retrieval principle: collect repeated failures as artifacts, then look across them for the pattern. The mining-your-transcripts move also maps onto how this vault should treat its own log.md.

The soft spot is verifiability itself. Her whole trust story leans on domains where the agent can run the real thing and prove it works. Verification is strong for front-end and type-level code, and thin exactly where the host pushes: finance, medicine, migrations that drop data. She concedes it on air. The page records the concession, not the sales pitch. The number “2,500 PRs in a month” is also softer than it sounds by her own account, since a lot of it is gardening and reading the code, not 2,500 features.