Outcome-driven, command-verified agent execution. Check what you can't trust.
Give Anvil an outcome, and it runs an agent in an isolated git worktree, looping
and feeding back each failure until a gate passes or it hits the attempt cap. The
gate decides done: a command you supply with --verify, or an auto-detected
build, typecheck, and test.
Anvil is the reliability-first spine of forge,
extracted clean: the gate and the loop are the most-tested, most paranoid code in
the repo, and orchestration is thin glue on top. See
docs/design.md for the contract and scope lock.
- Node.js >= 22.19.0
- A model API key. The default models route through Vercel's AI Gateway, so set
AI_GATEWAY_API_KEY(or pass--model <provider>:<id>and set that provider's key, e.g.ANTHROPIC_API_KEY).
npm install -g @vieko/anvil # the `anvil` command, globally
# or run it without installing:
npx @vieko/anvil run "implement feature X" -C /path/to/repogit clone https://github.com/vieko/anvil.git
cd anvil
npm install
npm run build
npm link --workspace @vieko/anvil # makes `anvil` available globallyOr run it straight from source without linking:
npm run dev -- run "implement feature X" -C /path/to/repoanvil run "<outcome>" # outcome to its gate, in an isolated worktree
anvil run specs/feature.md # an outcome from a spec file
anvil run "..." -C /path/to/repo # target another repo (like git -C)
anvil run "..." --verify "npm test" # explicit gate (repeatable; all must pass)
anvil run "..." --contract test/parser.test.ts # seed a frozen test the agent must satisfy
anvil run "..." --scope "src/**" # fence the agent into these paths
anvil run "..." --json # machine-readable result on stdout
anvil status # recorded runs: state, tokens, cost, and a spend total
anvil status --all --since 7d # the week's spend across every repo anvil ran in
anvil skills get core # print the full agent usage guideKey options (anvil --help for the rest):
| Option | Purpose |
|---|---|
-C, --dir <repo> |
Target repository (default: cwd). |
--from <ref> |
Fork the worktree from this ref (default: HEAD); e.g. main while on a feature branch. |
--verify "<cmd>" |
Gate command, repeatable. Omit it and Anvil auto-detects typecheck/build/test from package.json. |
--gate-crash-pattern <regex> |
Extra gate-output pattern identifying a verifier harness crash (repeatable). |
--no-baseline |
Skip the untouched-fork gate run before the first dispatch. |
--contract <file> |
Seed a check (typically a failing test) into the worktree and freeze it: the agent must satisfy it, never edit it. The strongest gate. |
--scope <glob> |
Fence the agent into these paths; a change outside voids the run. |
--model <alias|provider:id> |
Base model: haiku / sonnet / opus / fable / astra / sol / luna / terra / glm, or a concrete provider:model-id. Default sonnet. |
--effort <level> |
Base reasoning effort: low / medium / high / xhigh / max. Default high. |
-n, --max-attempts <n> |
Attempt cap before giving up (default 3). |
-v, --verbose |
Stream the agent's tool calls + gate progress to stderr. |
--reasoning |
Also stream the agent's reasoning trace (implies -v). Display-only — shows the thinking the model emits; use --effort to set its level. |
--json |
Emit a machine-readable result; human chrome and -v move to stderr. |
- Anvil cuts an isolated linked worktree on a fresh branch
anvil/<id>/<ts>. Your working tree is never touched, so a run can't stomp uncommitted work. - The agent works the outcome inside that worktree (
read/edit/write/bashtools). - Before any work, the gate checks the untouched fork SHA (unless
--no-baseline): a green baseline warns that the gate proves nothing unless the work adds checks; a red baseline is allowed and the agent proceeds. - The gate runs your verification commands in a clean environment and has
the only vote on "done":
- pass → Anvil commits the work and stops.
- fail → the errors feed the next attempt, and the config escalates:
one same-model retry a notch up in effort (the errors alone fix most
cheap-model failures), then the strong tier (
sonnetby default, up toopusonly when the gate keeps failing). - inconclusive (a flake, a timeout, no gate, or an identified verifier harness crash such as its own script being missing) → Anvil re-runs the gate in place (same attempt, same model, no re-dispatch) rather than call it a pass or a failure; a gate still inconclusive after 2 re-runs voids the run instead of paying a stronger model to face it. Failures in code under test remain ordinary failures.
- The loop ends at the attempt cap. State persists outside your repo, so
statusis exact and a passed outcome is never redone.
The result lives on its branch. Review, then merge:
git -C <repo> merge anvil/<id>/<ts> # or cherry-pick the commitAn agent that can edit the check can pass anything. The guards stop that:
--contract <file>: seed the check the agent must clear, out of its reach. Edit it, and the run is void.--scope <glob>: fence the agent into a set of paths. A change outside voids the run.- false-pass guard: an empty or provider-errored turn never reaches the gate.
All three can only force a no. The gate stays the only path to "done".
For a script or another agent driving Anvil, anvil run --json emits one object
(and anvil status --json the record ledger). Past the verdict it carries gate
provenance, how strong the green is, so the caller can route trust by rule
instead of re-reading the diff:
timeline is the per-attempt history (#12 Tier 3): what each attempt actually
dispatched, its verdict, and its own usage -- not just the final tally. Each
usage.cost is that attempt's USD spend (unrounded; absent when the model has no
price table), and the top-level usage is the cumulative sum; the human verdict
line shows the same total as $4.21. Cost is priced by anvil from the resolved
model's table: with PI_CACHE_RETENTION=long every Anthropic cache write is a 1h
write billed at 2x input, which pi's own usage.cost misses through the Vercel AI
Gateway (earendil-works/pi#9210).
anvil status appends each run's context tokens (2.3M ctx) and cost ($4.21)
plus a footer (N runs, P passed, F failed, $X.XX); --since 7d|24h|90m|<ISO>
filters by updatedAt, and --all reads every repo bucket under the state root
with each row prefixed by its repo name.
Anvil is a standalone CLI, so any agent or CI script can drive it: phrase the
outcome, call anvil run --json, and act on the verdict. No plugin, no lock-in
to one harness; Claude Code, Codex, and
the rest all work. anvil skills get core prints a guide an agent reads to learn
the conventions.
One stream can't serve a human and an agent at once, so don't ... -v 2>&1 | tee
(the pipe drops the TTY, output goes plain, and ANSI would pollute the log
anyway). Separate the channels instead — --json puts the clean result on
stdout and moves the colored -v stream + header to stderr:
# agent reads the JSON file; the human watches a colored -v stream in the pane
anvil run spec.md --verify "npm test" -v --json > result.json
echo "exit=$?" # result.json also carries { "passed": ... }If you genuinely need color through a pipe (a CI log, less -R, or a terminal
that misreports isTTY), set FORCE_COLOR=1; FORCE_COLOR=0 or NO_COLOR
force plain (NO_COLOR always wins).
Built on Pi.
npm install
npm run check # the one gate: biome -> tsc --noEmit -> build -> test
npm test # vitest only@anvil/core is runtime-agnostic; its node-bound seams (worktree, gate, agent)
live behind @anvil/core/node. Those seams (Agent, Workspace, Gate,
StatePersister) are injected interfaces, so tests drive the loop with fakes,
no real model, git, or filesystem needed.
Animation by Jon Romero Ruiz.

{ "id": "parser-tests", "passed": true, "attempts": 2, "timeline": [ { "attempt": 0, "config": { "model": "sonnet", "effort": "high" }, "verdict": "retrying", "usage": { "input": 812, "output": 340, "cacheRead": 0, "cacheWrite": 0, "cost": 0.0087 }, "errors": "...", "startedAt": "...", "endedAt": "..." }, { "attempt": 1, "config": { "model": "opus", "effort": "high" }, "verdict": "passed", "usage": { "input": 1204, "output": 512, "cacheRead": 6300, "cacheWrite": 0, "cost": 0.0553 }, "startedAt": "...", "endedAt": "..." } ], "usage": { "input": 2016, "output": 852, "cacheRead": 6300, "cacheWrite": 0, "cost": 0.064 }, "finalModel": "opus", "finalEffort": "high", "branch": "anvil/parser-tests/lz4k9", "gate": { "commands": ["tsc --noEmit", "npm test"], "source": "explicit" }, "contract": true, // a contract was enforced (and, since a violation voids the run, held) "scope": true // a scope was enforced (and held) }