Planned posts
- Sandboxes: containers, worktrees, egress allowlists and a credential proxy, so an agent can run without permission prompts.
- Pipelines: phases, typed hand-offs, tool boundaries per step, and why code owns the sequence.
- Gates: accepting work on evidence the agent didn’t write, and sending failures back to the right session.
- Budgets and metrics: spending limits enforced by the spender, pricing transcripts correctly, and cost per surviving line of code.
- Memory: pushed facts, pulled knowledge, and why nothing becomes memory without a person confirming it.
- Skills: scoping, sourcing and versioning skills across a pipeline.
- Harnesses: running Claude Code, Codex, Copilot CLI, Junie CLI and others behind one contract, and comparing them fairly.
- Review: a review tool built for agent output, in the IDE, the terminal or the browser: files grouped by kind, comments with a scope and a kind, and a review that goes back to the agent as its next brief.
- Fleets of agents: a planner that splits the work, checks on the division before a fleet starts, and a board for agents in separate sandboxes to talk on.
- The feedback loop: judging finished runs and improving the factory without letting it rewrite its own rules.
These follow one by one after the introduction above, and the order may shift as they get written.