A Claude Code plugin for taking work from concept to delivery, with you at three gates.
Five roles carry the work, a discipline for each part of it, and three times the run stops so you can decide whether it continues. It is plain markdown, plus one vendored bash library that walkthrough builds throwaway scripts from. Nothing here needs code running to stay alive.
This repository's own glossary defines every word it uses (see document home if you've pointed yours elsewhere), and DESIGN.md holds the reasoning behind the shape.
claude plugin marketplace add BytesNation/capstan
claude plugin install capstan@bytesnationRestart Claude Code once it finishes. Agent definitions load at session start, so nothing applies until you do.
That installs at user scope, so the crew is available in every session. Add --scope project to write to the project's .claude/settings.json instead, which is what you want when a team shares one repository.
Run /capstan:setup next, in the repository you plan to work in. It asks where the glossary, decision log, decision records and tracker should live, in the repository by default or in a vault, and writes the answer down so nothing asks again. Skip it and the first effort you start asks the same question.
/capstan:effort add rate limiting to the public API
You are now talking to the Architect, and you keep talking to it for the rest of the run. It owns the interview, the spec, the slice graph and the decision log. It does not build and it does not review, because the Builder and the Reviewer do that, and keeping those apart is the whole reason there are five roles instead of one.
What happens, in order:
- It interviews you. Rounds of questions, each carrying a recommended answer, so most rounds are you confirming rather than composing. Finding things out is its job, not yours. Anything the codebase or a primary source can answer, it goes and reads.
- It stops at gate one with a brief covering what it is building, why, and what it is deliberately not building. Protect this gate. A wrong turn is cheapest to catch here, and it is the one that gets waved through because the concept feels obvious to everyone in the room.
- Phase two plans. It cuts the work into vertical slices, agrees where the tests go, and stops at gate two with the slice graph and everything it had to assume.
- Phase three builds. One Builder per slice, each in its own git worktree, each writing a failing test first. A Reviewer reads every diff without the Builder's reasoning, on two axes that never blend into one verdict. Once the slices merge, your repository's own checks run against the integration, because passing alone proves nothing about passing together. Gate three shows you what was built, what review found, and what verification showed.
- Phase four delivers. The Courier packages the output and writes the permanent note. You commit it, once the note review passes. It never sends anything to anyone. That part stays yours.
| Gate | The brief answers | You decide |
|---|---|---|
| 1. Concept locked | What we are building, why, what we are explicitly not doing | Right thing? |
| 2. Plan locked | How, cut into slices, what runs parallel, what was assumed | Right shape? |
| 3. Ready to deliver | What was built, what review and verification found, what goes to whom | Ship? |
The run ends at every gate. Nothing polls, nothing waits in the background, and the crew never asks whether it should stop. The run being over is what makes the gate real. .capstan/effort/CLAIM.md records which phase the effort reached and what is still outstanding inside it, so the next session picks up where the last one stopped, however many hours later.
Unclear requirements never stop the run. The crew takes the most defensible reading, writes the assumption into the decision log, and keeps going. Every assumption surfaces at the next gate, where correcting one costs almost nothing. Four things do stop it: secrets and credentials, anything a third party will see, anything that costs money, and anything destructive or production-facing.
Three efforts at once is the ceiling. Three gates each against one reader is nine briefs a cycle, which is about where briefs stop being read and start being rubber-stamped. Fan-out inside a single effort has no limit.
The disciplines the roles pull in, plus the two front doors: effort starts a run, and setup configures where its artifacts live. Both are invoked only by you, and neither is a discipline.
| Skill | For |
|---|---|
interview |
Rounds of questions, each carrying a recommended answer, so decisions stay yours. |
spike |
A throwaway build, so a stalled design question gets something concrete to react to. |
slicing |
Cuts a locked plan into vertical slices with real blocking edges. |
test-first |
Red, green, refactor, tested only at seams agreed in advance. |
diagnosing-bugs |
A feedback loop that goes red on the bug before anyone theorises about the cause. |
codebase-design |
The words for structure, so a review can say a module is too shallow instead of that it feels wrong. |
two-axis-review |
Standards and spec, reviewed independently, never blended into one verdict. |
verify |
Runs the checks your repository declares against the merged result, because a slice passing alone proves nothing about the ones beside it. |
resolving-merge-conflicts |
Integrating parallel Builders, where neither side of a conflict can be asked what it meant. |
walkthrough |
The one-time script that carries you through a manual procedure, stage by stage, capturing what comes back. |
decision-record |
A one-line log by default, a full record only when one is earned. |
brief |
Checkpoint and partner briefs, generated per recipient rather than maintained. |
to-questionnaire |
Turns a question nobody in the room can answer into a document for the person who can. |
unslop |
Cuts AI tells from prose a person reads. |
writing-for-agents |
Keeps a document an agent consumes flat and the same shape every run. |
By default, everything lands in .capstan/, inside the repository the work is happening in:
.capstan/
CONTEXT.md one line per term. committed. edited in place.
decisions.md one line per decision. committed.
decisions/ a full record, only when one is earned. committed.
tracker.md one row per slice: effort, slice, status, merge commit. committed.
effort/ scratch: the claim, spec, plan, scout returns. gitignored.
effort and setup both make sure .capstan/effort/ is in your .gitignore, so you never have to add it by hand. That folder is deleted at delivery, because a stale spec is worse than no spec: the next agent reads it as current.
tracker.md answers what shipped, one row per slice, with the commit that merged it, and it is on by default, so you get it without configuring anything.
Capstan reads two more things from your own CLAUDE.md or AGENTS.md rather than hardcoding them. The permanent per-effort note's destination is yours to set, and the Courier skips the note and says so if you have not set it. The document home described below also lives in one of those files, but setup writes that line for you rather than you writing it yourself.
By default, the glossary, the decision log, the decision records, and the tracker live in .capstan/ in the repository, next to your code. That is the layout above, and the first effort you run writes it down as soon as you answer its one question with the default, rather than leaving anything unconfigured.
The effort scratch never moves, whichever layout you pick below: the claim, the spec, the plan, and scout returns stay at .capstan/effort/ in the repository under every configuration, and are deleted at delivery.
Some operators would rather keep a growing decision log out of the repository entirely, or keep one project's notes fully apart from another's. Three layouts cover that. setup asks the fork first, here in the repository or somewhere outside it, and only asks which of the two vault layouts you want if you choose outside:
- in the repository: the default above, nothing to configure, and the files sit next to the code they document.
- one vault per project: a dedicated vault for this project alone, for someone who wants it kept apart from every other project.
- one folder per project in a shared vault: one vault holding every project as its own folder, for someone who wants related projects visible together.
Capstan stores no difference between the last two. Both are an absolute path configured away from the default, and the list above exists to help you choose, not because Capstan branches on which one it is.
Run /capstan:setup to choose, and run it again later to change your mind. It asks the fork first and the vault layout only if you go outside the repository, confirms the path, creates the folder if it does not exist, checks whether anything is already sitting at the destination before it moves a thing, and moves the four artifacts on your approval. It also writes capstan-document-home into whichever of CLAUDE.md or AGENTS.md you already have, creating AGENTS.md if you have neither.
Capstan never commits a configured document home, and it never runs git inside one. It writes the files and stops; committing them from then on is yours, the same as committing anything else in that vault. Left uncommitted, a document home can sit that way for days before anyone notices, so weigh that before you switch it on. A document home inside a vault synced by iCloud, Dropbox, or Obsidian Sync is untested.
capstan-document-home lives in this repository's own CLAUDE.md or AGENTS.md, the one at the root of this working copy, not a user-level file. A ~/.claude/CLAUDE.md is never read for this key: set it only there and Capstan falls back to the default without telling you.
capstan-document-home: /Users/you/vault/YourProjectThe value must be an absolute path. A file lives in exactly one location, never a copy in both places, and Capstan resolves every path to it against that one root.
The glossary (CONTEXT.md), the decision log (decisions.md), the decision records (decisions/), and the tracker (tracker.md) resolve there from then on. Switching an existing project's document home runs through setup, which reports what it finds at the new destination and moves the four artifacts once you approve, rather than leaving you with a stale copy in .capstan/ and a fresh one in the vault.
An unreachable configured root stops the run rather than falling back to the default, since a missing source of truth would otherwise produce two records that quietly disagree.
claude plugin marketplace update bytesnation
claude plugin update capstan@bytesnationThe first refreshes the marketplace catalogue and installs nothing by itself. The second does the install and needs the marketplace-qualified name: capstan@bytesnation resolves, a bare capstan does not. Add -s project if that is where you installed, then restart Claude Code.
Leave the /plugin auto-update toggle off, the same as for any third-party marketplace.
No command edits a marketplace's source in place, and editing ~/.claude/plugins/known_marketplaces.json by hand does not hold, because the next claude plugin marketplace update puts the old address back. Removing and re-adding is the route that sticks.
Removing a marketplace uninstalls every plugin installed from it. The install record empties and the marketplace clone is deleted, so the second and third commands here are the repair rather than tidying afterwards. Run the first alone and you have no Capstan:
claude plugin marketplace remove bytesnation
claude plugin marketplace add BytesNation/capstan
claude plugin install capstan@bytesnation --scope userUse --scope project on the last command if that is where it was installed. capstan@bytesnation survives the round trip because a marketplace takes its name from the name field in its marketplace.json rather than from the repository path, so re-adding from a different address produces the same marketplace and the same plugin identifier.
The version cache under ~/.claude/plugins/cache/ is untouched throughout. A session open while you do this keeps resolving skills from the copy it already holds, and picks up the reinstall when you restart it.
git clone https://github.com/BytesNation/capstan.git
cp -r capstan/agents/* ~/.claude/agents/
cp -r capstan/skills/* ~/.claude/skills/Run /setup next, in the repository you plan to work in, to choose where the glossary, decision log, decision records and tracker should live. Skip it and the first effort you start asks the same question. Then start work with /effort <what you want built>.
A plugin namespaces what it ships, so the Builder is capstan:builder under a plugin install and plain builder under a manual one. skills/effort/SKILL.md tells the Architect which agent to spawn for each role, so use whichever form your install produced. The plugin path is the exercised one: installing, updating, and a full remove and re-add have all been run against it. The manual copy is there for a setup with no marketplace, and sees less use.
Copying over an existing manual install leaves behind any file a new version dropped or renamed. Replace a skill outright rather than copying onto it:
rm -rf ~/.claude/skills/effort
cp -r capstan/skills/effort ~/.claude/skills/Most skills here are a lone SKILL.md. The ones below are not, and lifting only the SKILL.md out of one of them leaves pointers aimed at files that are not there.
skills/effort/
SKILL.md identity, the crew, the gates, the precondition, phase 1
PHASE-2-PLAN.md
PHASE-3-BUILD.md
PHASE-4-DELIVER.md
skills/writing-for-agents/
SKILL.md
SKILL-MECHANICS.md frontmatter, invocation, router skills
AUDIT.md the editing pass to run against a target document
LICENSE, CREDIT.md upstream is MIT, see the licence section
skills/walkthrough/
SKILL.md identity, how to author a stage, the two guards before a write leaves the machine
template.sh vendored library, never edited
LICENSE, CREDIT.md upstream is MIT, see the licence section
skills/codebase-design/
SKILL.md the vocabulary and its principles
DEEPENING.md dependency categories, seam discipline, replace-don't-layer testing
DESIGN-IT-TWICE.md parallel sub-agents designing one interface several ways
LICENSE, CREDIT.md upstream is MIT, see the licence section
skills/diagnosing-bugs/
SKILL.md vendored, three repointed lines
LICENSE, CREDIT.md upstream is MIT, see the licence section
skills/to-questionnaire/
SKILL.md vendored, two local changes
LICENSE, CREDIT.md upstream is MIT, see the licence section
skills/resolving-merge-conflicts/
SKILL.md vendored, two local changes
LICENSE, CREDIT.md upstream is MIT, see the licence section
skills/unslop/
SKILL.md vendored, two local changes
LICENSE, CREDIT.md upstream is MIT, see the licence section
The Architect reads the file for the phase it is in, so a run that reaches gate two with no PHASE-2-PLAN.md beside it has nothing to follow and improvises a plan phase instead. Take the whole directory.
/capstan:effort cannot be invoked by a model. It carries disable-model-invocation: true, so only you start an effort. An agent that can start work on its own authority can commit you to work you never asked for. It does mean kickoff is always something you type.
Bash is an escape hatch. Scout's read-only guarantee is structural: its tool grant holds no write tools, so the harness enforces it. Builder, Reviewer and Courier all hold Bash, so their "never do X" rules are prose rather than enforcement. Add a deny list to settings.json if you want the gates enforced instead of requested.
Builder runs with acceptEdits. File writes never prompt. Bash commands still can, which is where unattended fan-out tends to stall.
Effort is not supported on Haiku. Drop the effort: frontmatter line from any agent you point at a Haiku model.
Fan-out does nothing for single-artifact work. Parallel Builders need slices that own different files. A document, a video script, a single config file: each is one artifact and inherently one Builder. Software usually fans out because slices own different things. Most other work does not, and a one-slice plan there is correct rather than a failure to parallelise.
The knowledge-base note is always reviewed whole. The note is never committed, so it has no fixed point to diff against. The Reviewer reads the full note, and review is never skipped.
MIT. See LICENSE. Take it, change it, ship it.
Some skills here are not ours. Every one is MIT, and every one is redistributed with its own licence and a CREDIT.md in its folder recording exactly what changed:
skills/writing-for-agents/:SKILL.mdandSKILL-MECHANICS.mdby Matt Pocock.AUDIT.mdbeside them is ours.skills/unslop/:SKILL.mdby Lauren Tan, via cursor/plugins. Two lines changed.skills/walkthrough/:template.shby Matt Pocock, vendored byte-identical.SKILL.mdbeside it is ours, written fresh around the library.skills/diagnosing-bugs/:SKILL.mdby Matt Pocock. Three lines repointed at Capstan's own paths and atwalkthrough.skills/codebase-design/:SKILL.md,DEEPENING.mdandDESIGN-IT-TWICE.mdby Matt Pocock. One line repointed; the other two files are byte-identical.skills/to-questionnaire/:SKILL.mdby Matt Pocock. Two changes: the invocation flag, and where the document lands and where its answers go.skills/resolving-merge-conflicts/:SKILL.mdby Matt Pocock. Two changes: where a hunk's intent is found, and one exception to never aborting.
Beyond those, nothing is vendored, though some ideas are borrowed, all from Matt Pocock. The prose is ours; the mechanics are his.
- The frontier in
interview: a design tree, where a question depending on an open question waits for a later round. Sharpened fromgrilling. - The
nextline inCLAIM.md: what the run after this one picks up, written for the agent that resumes rather than the person at the gate. Fromhandoff, sized down to a field in a file that already exists. - The
unformedstatus in the decision log: an area nobody can phrase a question about yet. His fog of war fromwayfinder, without the issue tracker it is charted on. - Two moves in
interview: challenging a term against the glossary rather than only within the session, and inventing an edge-case scenario when a relationship between concepts stays vague. Fromdomain-modeling, minus its file layout.