Your agent just deleted a file it wasn't supposed to touch.
You know the moment. You asked for a landing page tweak, it "noticed" your API calls were slow, and now two features are broken and you're scrolling the diff trying to figure out what else it improved for you.
Ras Mic doesn't have that problem. He runs a software factory: four steps, five or six markdown files, and up to 15 features being built in parallel by different agents without stepping on each other.
He walked Greg Isenberg through the whole system on the Startup Ideas podcast, screen share and all. The video did 117K views in four days, and the phrase "software factory" is suddenly everywhere.
I pulled out the full system so you can steal it. Watch the original here, then let's take it apart:
What a software factory actually is (and isn't)
Mic's definition, compressed: a software factory is your workflow, your skills, and your domain knowledge, packed into markdown files that tell an agent exactly how to build.
It is not a product, and it's not something you buy from one of the startups that slapped "software factory" on a landing page last month. (No shade, says Mic. But also, kind of shade.)
It's model agnostic and use agnostic. Claude Code, Codex, Cursor, whatever. The factory is the process, not the tool.
Think about how you built your last app. You typed "build me this", it built something, you didn't like it, you typed again. Repeat for six hours.
The factory replaces that back-and-forth with an assembly line. Each step squeezes the most out of the model, and the whole thing moves fast without shipping slop.
The centerpiece is the agents.md file, the document that gets injected into the chat before every single message you send. Mic's take: most people's agents.md is useless because it describes the codebase, which the agent can already read. His contains only the one thing the agent can't figure out on its own. The workflow.
Four steps: isolate, build, prove, ship.
Step 1: Isolate. Every feature gets its own station
First rule of the factory: never build on main.
Mic has a skill called "new feature". Every feature starts in a fresh git worktree branched from origin main. A worktree is basically a copy of the app at this exact moment. The agent works on the copy, then merges back when it's done.
Why this matters: when your agent "deleted a bunch of stuff", it's almost always because you had it working on two features on the same branch. You asked for a faster API. It saw the landing page calling the API "wrong" and rewrote it. The landing page you were also redesigning. In the same session.
The agent did what you told it to do. You just gave it a station where two orders collide.
With isolation, Mic showed four terminal tabs running at once: one agent building an email client, one building a Linux computer environment, one updating a landing page. On another project he had 15 features in flight simultaneously, each agent in its own worktree, with zero conflicts between them.
Greg's translation is the one that stuck with me: no real engineering team pushes straight to main with everyone editing the same files. Your agents are a team now. Structure them like one.
Step 2: Build. Force a code structure, or you'll get slop that works
Models will happily write bad code that runs. That one took me a while to accept.
Mic put it bluntly: if the model can get it done in a sloppy way, it'll get it done in a sloppy way. He's had GPT 5.6 Soul write a feature that worked perfectly, then had Fable review it. Verdict: duplicated functions, dead code, logic scattered everywhere. His quote: "This is disgusting."
It works. It's still a liability. The day you hire a developer (or another agent with no context) to touch that code, you pay the bill.
So the second skill is "code structure". It makes every agent write in a service layer architecture: a boring, readable pattern where a human developer, or a fresh agent, can open the codebase and immediately know where things live.
You don't need to know what service layer architecture is to use this. The point is: pick a structure, write it into a skill, and make every agent follow it. The structure is your factory's assembly spec. This is the same idea behind Claude skills: package the knowledge once, and every agent run inherits it.
Step 3: Prove. Agents can't pinky promise
My favorite step, because it kills the trust problem.
Agents lie. Not maliciously, but push one hard enough and it'll tell you the work is done when it isn't, or "realize" mid-conversation that it never actually made the fix.
Mic's rule: the agent has to prove its work with evidence, not claims.
His "evidence-driven testing" skill records a before state and an after state. Fixing a bug? Record the bug happening, then record the fixed version working. Machine can't record video? A fallback skill takes before and after screenshots instead.
Every PR his agents open comes with that proof embedded in the description. New admin email page: screenshot of the page not existing, screenshot of it working. Performance fix: one page was loading in 815 milliseconds ("a sin in web development"), the agent tuned it down to 61ms and posted the test numbers as proof.
And when the proof isn't there, the skill sends the agent back to build on its own, without you having to nag it.
The result: Mic barely reads code anymore. He flips through before-and-after proof the way you flip through Instagram stories. Greg called it exactly that, and Mic didn't argue.
Step 4: Ship. The Greptile loop that reviews the reviewer
Last station: quality control by someone who isn't the builder.
Mic wires in Greptile, a code review agent, through a skill he calls Grep Loop. (CodeRabbit and Macroscope do the same job. Most have free tiers you can run for a long time without paying.)
Every PR gets reviewed and scored with a confidence rating out of 5. Watch what happens when the score comes back low:
- Greptile scores the PR a 3/5 and lists what's wrong (a pagination bug, a menu state that didn't persist)
- The building agent reads the feedback and goes back to step 2: build
- Then step 3: prove the fix
- Then ships again and waits for a new score
- 4/5. One thing left. Loop again.
- 5/5. Now, and only now, does Mic enter the picture.
At 5/5 his entire job is clicking merge. The worktree gets cleaned up, the feature lands on main, and three other features are still moving through the line behind it.
Nobody tells the agent to run the loop. The skill defines the standard, and the agent keeps working until the standard is met. If you've read my breakdown of the AI agent loop, this is that idea taken to its logical end: the loop doesn't close until an external reviewer signs off.
Why this matters if you're a solo founder
Because this is how one person ships like a team of eight.
Every bootstrapped founder I talk to is somewhere on this curve. Stage one: chatting with Cursor, praying. Stage two: agents doing real work, but you review everything because you got burned. Stage three: systems doing the reviewing, and you managing the system.
The factory is stage three, and the barrier is embarrassingly low. It's markdown files. Mic gives his away for free (link in the video description), and you don't need another subscription or another platform to run them.
Your starter kit, condensed:
→ An agents.md that contains your workflow, not a description of your codebase
→ A "new feature" skill: fresh worktree per feature, never build on main
→ A "code structure" skill: one architecture pattern, enforced everywhere
→ A "prove it" skill: before and after evidence on every PR
→ A code review agent with a score threshold the builder has to hit
One warning from Mic worth repeating: don't blindly copy his files. The factory works because it encodes his workflow and his domain knowledge. Copy the structure, then fill it with how you work.
And the concept travels beyond code. The pattern (isolate the task, define the standard, demand proof, loop until an external check passes) is just good systems thinking. It's the same shape as the business automations that run your ops while you sleep. Define the standard once, let the machine hold itself to it.
The line from the episode that I can't stop thinking about: "Some startups are now an agent with a couple markdown files."
We're early. A few markdown files and a review loop put you ahead of most funded engineering teams right now, and that window won't stay open forever.
FAQ
What is a software factory in AI development?
A software factory is a repeatable workflow that lets AI agents ship production software in parallel without human babysitting. Ras Mic's version has four steps (isolate, build, prove, ship) defined in five or six markdown skill files, so it works with any model and any coding use. It's a method, not a product.
Do I need a specific AI model or tool to build one?
No. The factory is model and use agnostic by design. Mic runs it across different models (he names GPT-6 Astra as his workhorse and praises Fable's code quality) and the skills are plain markdown, so they drop into Claude Code, Codex, Cursor, or whatever you already use.
Do I really need a code review agent like Greptile?
If real users will touch your software, yes. The review agent is the only step where someone other than the builder checks the work, and the confidence score is what powers the automatic fix loop. Greptile, CodeRabbit, and Macroscope all work, and their free tiers are enough to start.
Can a non-technical founder run a software factory?
That's who it helps most. The prove step turns code review into flipping through before-and-after screenshots and videos, so you can verify work without reading a line of code. Greg Isenberg, who calls himself non-technical, frames the whole system as managing agents like an engineering team.
Want to see how real founders actually ship?
Every week on the Profitable Founder Podcast I sit down with bootstrapped founders doing $100K to $10M a year and pull out the real systems: the workflows, the numbers, what broke along the way.