Skip to content

Repository files navigation

Anvil

Anvil

Outcome-driven, command-verified agent execution. Check what you can't trust.

Give Anvil an outcome, and it runs an agent in an isolated git worktree, looping and feeding back each failure until a gate passes or it hits the attempt cap. The gate decides done: a command you supply with --verify, or an auto-detected build, typecheck, and test.

Anvil is the reliability-first spine of forge, extracted clean: the gate and the loop are the most-tested, most paranoid code in the repo, and orchestration is thin glue on top. See docs/design.md for the contract and scope lock.

Requirements

  • Node.js >= 22.19.0
  • A model API key. The default models route through Vercel's AI Gateway, so set AI_GATEWAY_API_KEY (or pass --model <provider>:<id> and set that provider's key, e.g. ANTHROPIC_API_KEY).

Install

npm install -g @vieko/anvil          # the `anvil` command, globally
# or run it without installing:
npx @vieko/anvil run "implement feature X" -C /path/to/repo

From source (for development)

git clone https://github.com/vieko/anvil.git
cd anvil
npm install
npm run build
npm link --workspace @vieko/anvil   # makes `anvil` available globally

Or run it straight from source without linking:

npm run dev -- run "implement feature X" -C /path/to/repo

Usage

anvil run "<outcome>"                          # outcome to its gate, in an isolated worktree
anvil run specs/feature.md                     # an outcome from a spec file
anvil run "..." -C /path/to/repo               # target another repo (like git -C)
anvil run "..." --verify "npm test"            # explicit gate (repeatable; all must pass)
anvil run "..." --contract test/parser.test.ts # seed a frozen test the agent must satisfy
anvil run "..." --scope "src/**"               # fence the agent into these paths
anvil run "..." --json                         # machine-readable result on stdout
anvil status                                   # recorded runs: state, tokens, cost, and a spend total
anvil status --all --since 7d                  # the week's spend across every repo anvil ran in
anvil skills get core                          # print the full agent usage guide

Key options (anvil --help for the rest):

Option Purpose
-C, --dir <repo> Target repository (default: cwd).
--from <ref> Fork the worktree from this ref (default: HEAD); e.g. main while on a feature branch.
--verify "<cmd>" Gate command, repeatable. Omit it and Anvil auto-detects typecheck/build/test from package.json.
--gate-crash-pattern <regex> Extra gate-output pattern identifying a verifier harness crash (repeatable).
--no-baseline Skip the untouched-fork gate run before the first dispatch.
--contract <file> Seed a check (typically a failing test) into the worktree and freeze it: the agent must satisfy it, never edit it. The strongest gate.
--scope <glob> Fence the agent into these paths; a change outside voids the run.
--model <alias|provider:id> Base model: haiku / sonnet / opus / fable / astra / sol / luna / terra / glm, or a concrete provider:model-id. Default sonnet.
--effort <level> Base reasoning effort: low / medium / high / xhigh / max. Default high.
-n, --max-attempts <n> Attempt cap before giving up (default 3).
-v, --verbose Stream the agent's tool calls + gate progress to stderr.
--reasoning Also stream the agent's reasoning trace (implies -v). Display-only — shows the thinking the model emits; use --effort to set its level.
--json Emit a machine-readable result; human chrome and -v move to stderr.

How it works

  1. Anvil cuts an isolated linked worktree on a fresh branch anvil/<id>/<ts>. Your working tree is never touched, so a run can't stomp uncommitted work.
  2. The agent works the outcome inside that worktree (read / edit / write / bash tools).
  3. Before any work, the gate checks the untouched fork SHA (unless --no-baseline): a green baseline warns that the gate proves nothing unless the work adds checks; a red baseline is allowed and the agent proceeds.
  4. The gate runs your verification commands in a clean environment and has the only vote on "done":
    • pass → Anvil commits the work and stops.
    • fail → the errors feed the next attempt, and the config escalates: one same-model retry a notch up in effort (the errors alone fix most cheap-model failures), then the strong tier (sonnet by default, up to opus only when the gate keeps failing).
    • inconclusive (a flake, a timeout, no gate, or an identified verifier harness crash such as its own script being missing) → Anvil re-runs the gate in place (same attempt, same model, no re-dispatch) rather than call it a pass or a failure; a gate still inconclusive after 2 re-runs voids the run instead of paying a stronger model to face it. Failures in code under test remain ordinary failures.
  5. The loop ends at the attempt cap. State persists outside your repo, so status is exact and a passed outcome is never redone.

The result lives on its branch. Review, then merge:

git -C <repo> merge anvil/<id>/<ts>   # or cherry-pick the commit

Guards

An agent that can edit the check can pass anything. The guards stop that:

  • --contract <file>: seed the check the agent must clear, out of its reach. Edit it, and the run is void.
  • --scope <glob>: fence the agent into a set of paths. A change outside voids the run.
  • false-pass guard: an empty or provider-errored turn never reaches the gate.

All three can only force a no. The gate stays the only path to "done".

Machine-readable output

For a script or another agent driving Anvil, anvil run --json emits one object (and anvil status --json the record ledger). Past the verdict it carries gate provenance, how strong the green is, so the caller can route trust by rule instead of re-reading the diff:

{
  "id": "parser-tests",
  "passed": true,
  "attempts": 2,
  "timeline": [
    { "attempt": 0, "config": { "model": "sonnet", "effort": "high" }, "verdict": "retrying", "usage": { "input": 812, "output": 340, "cacheRead": 0, "cacheWrite": 0, "cost": 0.0087 }, "errors": "...", "startedAt": "...", "endedAt": "..." },
    { "attempt": 1, "config": { "model": "opus", "effort": "high" }, "verdict": "passed", "usage": { "input": 1204, "output": 512, "cacheRead": 6300, "cacheWrite": 0, "cost": 0.0553 }, "startedAt": "...", "endedAt": "..." }
  ],
  "usage": { "input": 2016, "output": 852, "cacheRead": 6300, "cacheWrite": 0, "cost": 0.064 },
  "finalModel": "opus",
  "finalEffort": "high",
  "branch": "anvil/parser-tests/lz4k9",
  "gate": { "commands": ["tsc --noEmit", "npm test"], "source": "explicit" },
  "contract": true,  // a contract was enforced (and, since a violation voids the run, held)
  "scope": true      // a scope was enforced (and held)
}

timeline is the per-attempt history (#12 Tier 3): what each attempt actually dispatched, its verdict, and its own usage -- not just the final tally. Each usage.cost is that attempt's USD spend (unrounded; absent when the model has no price table), and the top-level usage is the cumulative sum; the human verdict line shows the same total as $4.21. Cost is priced by anvil from the resolved model's table: with PI_CACHE_RETENTION=long every Anthropic cache write is a 1h write billed at 2x input, which pi's own usage.cost misses through the Vercel AI Gateway (earendil-works/pi#9210).

anvil status appends each run's context tokens (2.3M ctx) and cost ($4.21) plus a footer (N runs, P passed, F failed, $X.XX); --since 7d|24h|90m|<ISO> filters by updatedAt, and --all reads every repo bucket under the state root with each row prefixed by its repo name.

Drive it from any harness

Anvil is a standalone CLI, so any agent or CI script can drive it: phrase the outcome, call anvil run --json, and act on the verdict. No plugin, no lock-in to one harness; Claude Code, Codex, and the rest all work. anvil skills get core prints a guide an agent reads to learn the conventions.

One stream can't serve a human and an agent at once, so don't ... -v 2>&1 | tee (the pipe drops the TTY, output goes plain, and ANSI would pollute the log anyway). Separate the channels instead — --json puts the clean result on stdout and moves the colored -v stream + header to stderr:

# agent reads the JSON file; the human watches a colored -v stream in the pane
anvil run spec.md --verify "npm test" -v --json > result.json
echo "exit=$?"   # result.json also carries { "passed": ... }

If you genuinely need color through a pipe (a CI log, less -R, or a terminal that misreports isTTY), set FORCE_COLOR=1; FORCE_COLOR=0 or NO_COLOR force plain (NO_COLOR always wins).

Built on Pi.

Develop

npm install
npm run check     # the one gate: biome -> tsc --noEmit -> build -> test
npm test          # vitest only

@anvil/core is runtime-agnostic; its node-bound seams (worktree, gate, agent) live behind @anvil/core/node. Those seams (Agent, Workspace, Gate, StatePersister) are injected interfaces, so tests drive the loop with fakes, no real model, git, or filesystem needed.

Credits

Animation by Jon Romero Ruiz.

License

MIT

About

Outcome-driven, command-verified agent execution. Check what you can't trust.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages