Start in 30 seconds · GitHub Action · Plugin demo · Plugin suites · MCP contracts · Agent teams · Doctor · Distribution · Security
Deterministic regression testing and compatibility intelligence for Claude Code.
Claude Code changes. Your CLAUDE.md, hooks, plugins, MCP servers and permission rules change too. Canary gives you a repeatable answer to the question that normally becomes guesswork:
What broke, and which Claude Code release first broke it?
Canary runs the same scenario from the same Git commit in disposable worktrees, captures tool/token/reported-cost/duration metrics, checks deterministic assertions, compares releases, bisects regressions and builds full plugin compatibility matrices.
Canary v2 turns those individual checks into a compatibility platform:
- first-class scenario suites with tags, affected-path selection, concurrency, deterministic sharding and run budgets;
- scheduled release watch with known-good state, regression detection and automatic first-bad-release bisection;
- deterministic failure fingerprints and flakiness analysis so noisy scenarios are not confused with upstream regressions;
- portable HTML, JUnit and SARIF reporting plus local historical trends;
- compatibility manifests,
canary.lock, open registry aggregation, evidence-backed badges and inspectable scenario packs; - permission-policy/trust regression coverage, isolated MCP fixtures, gateway matrices and signed/checksummed attestations;
- versioned public schemas plus compatibility query/explain/graph APIs for build tools and multi-project workspaces.
All existing v1 workflows remain available through the v2 CLI.
Install the published CLI and verify the host before spending tokens:
npm install -g claude-code-canary
claude-canary doctor
claude-canary init
claude-canary run .canary/basic.canary.ymlAlready using GitHub Actions? Add Canary as a CI step and use the stable @v2 major channel:
- uses: SLP-DEV1/claude-code-canary@v2
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
with:
mode: run
scenario: .canary/basic.canary.ymlThat gives you the same scenario locally and in CI, with machine-readable results in .canary/results/ and a Markdown summary in GitHub Actions.
Choose the workflow that matches the problem:
| Goal | Command / Action mode |
|---|---|
| Catch a Claude Code release regression | compare |
| Find the first bad Claude Code release | bisect |
| Gate a repository pull request | pr-check |
| Keep recurring CI to one Claude run | baseline-check |
| Test an entire Claude Code plugin surface | plugin-suite |
| Detect MCP schema/capability drift | mcp-check |
| Run a deterministic scenario suite | suite |
| Guard against newly published Claude Code releases | watch |
| Measure scenario stability/noise | flake |
| Publish/query portable compatibility evidence | compat / lock |
| Generate interoperable local/CI reports | report / trend |
| Check host/plugin/MCP readiness without exposing secrets | doctor |
Plugin author? Turn your plugin surface into smoke tests, then test every generated scenario across recent Claude Code releases:
claude-canary plugin-init ./my-plugin
claude-canary plugin-suite --plugin ./my-plugin --last 10You get one report like this:
| Claude Code | load | command-review | skill-api | hook-stop | mcp-github | Overall |
|-------------|:----:|:--------------:|:---------:|:---------:|:----------:|---------|
| 2.1.231 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ Compatible |
| 2.1.232 | ✅ | ✅ | ✅ | ❌ | ✅ | ❌ 1 failed |
| 2.1.233 | ✅ | ❌ | ✅ | ❌ | ✅ | ❌ 2 failed |
Need the exact regression boundary instead?
claude-canary bisect .canary/plugin-smoke.canary.yml \
--good 2.1.220 \
--bad 2.1.237| Problem | Canary |
|---|---|
| "The new Claude release feels worse" | Run the same deterministic scenario on both releases. |
| "Which release broke us?" | Binary-search the real published Claude Code release range. |
| "Does this plugin still work?" | Generate smoke tests and run a release × component compatibility suite. |
"Is my new CLAUDE.md actually better?" |
Run interleaved A/B configuration experiments. |
| "This worked yesterday" | Record a good real task and replay it from the exact original commit. |
| "How do I report this bug safely?" | Export a bounded, redacted reproduction bundle. |
| "Did agent usage blow up?" | Track tool calls, tokens, duration and reported cost. |
| "Did this PR make the agent worse?" | Compare base vs head with the same Claude executable and fail on configured deltas. |
| "Can CI do this without paying for two runs every time?" | Commit a known-good metric baseline and execute only the candidate. |
| "Did my MCP server silently change?" | Snapshot tools/prompts/resources and fail on removed tools, schema changes or capability regressions. |
| "Did Claude change how my agent team coordinates?" | Observe real interactive teammate/task lifecycle signals and compare privacy-safe structural snapshots. |
| "Is this host actually ready for my extensions?" | Run a secret-free Doctor preflight for provider mode, plugins, LSP binaries, project MCP transports and agent-team constraints. |
Canary is not another transcript viewer or generic model leaderboard. It is a regression layer for real Claude Code workflows.
The v2 Action supports compare, run, pr-check, baseline-check, mcp-check, plugin-matrix, plugin-suite, suite and watch through one Marketplace-ready action.yml. pr-check can also update one stable pull-request comment with the regression table when comment-pr: true is enabled.
A plugin compatibility gate can be as small as:
name: Claude Canary
on:
workflow_dispatch:
push:
branches: [main]
permissions:
contents: read
jobs:
plugin-compatibility:
runs-on: ubuntu-latest
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
steps:
- uses: actions/checkout@v7
with:
fetch-depth: 0
- uses: SLP-DEV1/claude-code-canary@v2
with:
mode: plugin-suite
plugin: ./my-plugin
last: 10The Action streams progress into the job log, writes the combined Markdown report to the GitHub Step Summary and uploads .canary/results/ as an artifact. Inputs are converted into a direct argument array rather than interpolated into a shell command.
Security: do not expose
ANTHROPIC_API_KEYor other secrets to untrusted fork code/scenarios. Prefer trusted branches,workflow_dispatch, or carefully designed PR workflows. See GitHub Action security.
See docs/GITHUB_ACTION.md for every input/output and more examples.
Requirements:
- Node.js 20+
- Git
- an authenticated Claude Code installation or API-key environment suitable for headless Claude Code
Install the published CLI from npm:
npm install -g claude-code-canary
claude-canary --version
claude-canary doctorFor repository development, clone this repo and run npm ci --ignore-scripts && npm run build.
Before an extension-heavy run, use the machine-readable compatibility preflight:
claude-canary doctor --json
claude-canary doctor --plugin ./my-plugin --jsonIt reports only non-secret configuration shape: provider mode, credential-presence booleans, plugin component types, LSP/stdio-MCP executable availability, project MCP transport types and agent-team/TTY warnings. API keys, OAuth tokens, base URLs, MCP URLs/headers and environment values are never emitted. See Extension compatibility doctor.
Then create and run a scenario inside the repository you want to test:
cd /path/to/project
claude-canary init
claude-canary run .canary/basic.canary.ymlCanary drives the Claude Code CLI rather than calling a model API directly. That means a Claude Code setup that already works through a compatible gateway can normally be exercised by Canary through the same CLI configuration and environment.
A useful preflight check is:
claude -p "Reply exactly LOCAL_OK"If that command reaches your configured gateway and succeeds, Canary can invoke the same claude executable from an isolated worktree:
claude-canary run .canary/basic.canary.ymlA manual end-to-end smoke test has successfully exercised this path on Windows:
Claude Code
→ Claude Code Router
→ llama.cpp
→ Qwen3.8-27B
→ Claude Code Canary
This is a community/custom deployment path, not part of Canary's guarantee that historical Claude Code releases behave identically across third-party gateways. Proxy translation, model behavior and gateway routing can add their own sources of variance.
Canary labels total_cost_usd as reported cost. With a proxy or local model, that value may be estimated, synthetic or otherwise unrelated to actual billing. Treat it as upstream accounting metadata unless your provider explicitly documents it as billable cost.
Inspect a stdio MCP server directly, without invoking a model:
claude-canary mcp-snapshot .canary/mcp/github.mcp.yml
# review + commit the generated baseline
claude-canary mcp-check .canary/mcp/github.mcp.yml --require-baselineCanary snapshots tools and JSON Schemas, prompts, resources, resource templates, capabilities and observed list_changed signals. Removed tools, schema changes and disabled capabilities are breaking by default; additions are reported without failing CI. Tool safety annotations can also be asserted without executing the tool. See MCP contract testing.
Agent teams are an experimental upstream surface and currently require a real interactive Claude Code session. Canary keeps that distinction explicit:
claude-canary team-run examples/agent-team.team.yml --version latestThe temporary observer records teammate names/types, task IDs/state, message counts, idle transitions and stop failures — not teammate prompts, message bodies, task descriptions or transcripts. Saved snapshots can be compared non-interactively:
claude-canary team-compare baseline-agent-team.json candidate-agent-team.jsonIf stdin/stdout is not a real TTY, team-run returns unsupported rather than silently measuring ordinary -p subagents. See Agent-team regression testing.
Run the same scenario against the base and head Git refs with one Claude executable:
claude-canary pr-check .canary/basic.canary.yml \
--base origin/main \
--head HEADThis catches repository changes that keep the final task green but increase tokens/cost/tool calls, introduce permission prompts, or change configured hook semantics. See Pull request regression checks.
claude-canary baseline update .canary/basic.canary.yml
# commit .canary/baselines/<scenario-name>.json
claude-canary baseline check .canary/basic.canary.ymlBaselines use the same regression thresholds while cutting recurring CI from two Claude runs to one. A SHA-256 of the scenario prevents stale snapshots from silently passing after the scenario changes. See Committed baselines.
claude-canary compare .canary/basic.canary.yml \
--from 2.1.220 \
--to latestCanary keeps historical native binaries in its own cache and never replaces your normal claude installation. Release manifests are checksum-verified; signed manifests are signature-verified where Anthropic publishes signatures.
compare can also fail on relative regressions even when both releases still produce the correct result: token growth, reported-cost growth, extra tool calls, new permission prompts/denials, or a changed hook sequence. See Efficiency and lifecycle regressions.
claude-canary bisect .canary/basic.canary.yml \
--good 2.1.220 \
--bad 2.1.237Only the releases needed by binary search are executed. Like git bisect, this assumes a monotonic good → bad transition.
claude-canary plugin-init ./my-pluginCanary discovers standard and manifest-defined plugin surfaces including:
- commands
- agents
- skills
- hooks
- MCP servers
- LSP servers
- background monitors (static contract only; never auto-started by Canary)
- plugin dependencies and version constraints (static contract only)
Generated suites are marker-protected, symlink-safe and intentionally reviewable. They are scaffolds: strengthen assertions for domain-specific behavior before treating them as proof.
claude-canary plugin-suite \
--plugin ./my-plugin \
--last 10The suite refuses stale discovery metadata and missing generated scenarios so an old or accidentally incomplete suite cannot silently turn green. A default run budget prevents accidental scenarios × releases explosions.
claude-canary plugin-matrix \
.canary/plugins/my-plugin/command-review.canary.yml \
--plugin ./my-plugin \
--from 2.1.220 \
--to 2.1.237Use plugin-suite for the broad gate and plugin-matrix for focused debugging.
claude-canary experiment .canary/basic.canary.yml \
--baseline-config .canary/variants/current \
--candidate-config .canary/variants/candidate \
--runs 5Variants can cover project instructions, settings, rules, hooks, MCP config and local plugins. Canary interleaves baseline/candidate runs and reports pass-rate and efficiency deltas. Variant trees containing symlinks are refused.
claude-canary record auth-fix \
--prompt "Fix the failing authentication test without changing the public API" \
--setup "npm ci" \
--verify "npm test"
# Run the real Claude task, then:
claude-canary save auth-fix
# Later:
claude-canary replay .canary/auth-fix.canary.ymlThe generated scenario records the exact starting Git commit and deterministic changed-file expectations without persisting raw environment values or a Claude transcript.
claude-canary repro .canary/results/failed.jsonRepro bundles use bounded fixture selection, deny credential/build/cache paths, refuse symlinks and binaries, redact common secret/path patterns and generate Linux/macOS plus PowerShell launchers. Review every bundle before publishing it. Generic redaction cannot understand project-specific confidentiality.
version: 1
name: fix-auth-regression
prompt: |
Fix the failing authentication test.
Do not change the public API.
setup:
commands:
- npm ci
claude:
executable: claude
permission_mode: dontAsk
timeout_seconds: 900
verify:
commands:
- npm test
expect:
changed_files:
allow:
- src/auth/**
- test/auth/**
require:
- src/auth/**
deny:
- package-lock.json
file_contains:
- path: src/auth/index.ts
text: authenticate
permissions:
max_prompts: 0
max_denied: 0
hooks:
sequence:
- PreToolUse
- PostToolUse
limits:
max_tool_calls: 100
max_total_tokens: 200000
max_cost_usd: 5
regressions:
max_total_tokens_increase_pct: 25
max_reported_cost_increase_pct: 20
max_tool_calls_increase_pct: 25
max_permission_prompts_increase: 0
require_same_hook_sequence: trueA run passes only when Claude exits successfully and every configured deterministic assertion/limit can be evaluated and passes. v1 fails closed on malformed/truncated stream-json and on cost limits when Claude does not report cost.
Every ordinary run starts from a clean tracked repository state and executes in a disposable detached Git worktree. Canary separates generated results from the tested worktree and cleans temporary worktrees/runtime copies after execution.
Additional v1 hardening includes:
- bounded subprocess output capture;
- fail-closed malformed protocol handling;
- validated Claude release platform IDs;
- bounded release downloads plus checksum/signature verification;
- symlink refusal for plugin suites and configuration variants;
- marker-protected destructive regeneration/repro operations;
- exact direct dependency versions plus a committed lockfile;
- SHA-pinned third-party Actions;
- CodeQL and cross-platform CI.
Canary scenarios themselves can contain setup/verification shell commands and Claude permission options. Treat scenario/config files as trusted code, especially in CI with credentials.
Read docs/SECURITY_MODEL.md and SECURITY.md before using Canary on untrusted contributions.
v1 also exposes a supported ESM library entry point:
import {
loadScenario,
runScenario,
runPluginSuite,
formatPluginSuiteMarkdown,
CANARY_VERSION,
} from 'claude-code-canary';The public entry point is dist/api.js / dist/api.d.ts. Internal files are intentionally not package exports. The core run artifact contract is documented by schemas/run-result.schema.json.
init Create a starter scenario
validate Validate scenario YAML without spending tokens
run Run one deterministic scenario
compare Compare two executables or releases
pr-check Compare one Claude executable across two Git refs
baseline Create/check committed known-good metric baselines
bisect Find the first bad executable/release
experiment A/B test Claude Code configuration variants
record / save Capture a successful real task as a scenario
replay Replay from the recorded starting commit
repro Create a privacy-first bug reproduction bundle
plugin-init Discover a plugin and generate smoke scenarios
plugin-matrix Test one plugin scenario across releases
plugin-suite Test the complete generated plugin surface across releases
versions Install/list/locate isolated Claude Code releases
doctor Check local prerequisites and repository readiness
Run claude-canary <command> --help for command-specific flags.
| Guide | What it covers |
|---|---|
| GitHub Action | Marketplace usage, modes, inputs, outputs and CI security |
| Pull request checks | Base-vs-head regression gates and optional stable PR comments |
| Committed baselines | One-run CI against reviewed known-good metrics |
| Plugin suites | Full release × plugin-surface matrices |
| Plugin smoke generator | Discovery rules and generated scenarios |
| Plugin compatibility | Focused plugin matrices and isolation |
| Version manager | Release cache, checksums and signature trust |
| Efficiency & lifecycle regressions | Relative token/cost/tool regressions, permission prompts and ordered hooks |
| Configuration experiments | A/B test layout and interpretation |
| Record & replay | Recording workflow and privacy model |
| Reproduction bundles | Safe bundle generation and publishing checklist |
| Security model | Trust boundaries and threat model |
| Reproducibility | Determinism guarantees and unavoidable variance |
| Releasing | v1 tags, Marketplace and release checklist |
| Distribution | npm, Marketplace and curated-list publication checklist |
| Roadmap | What is shipped and what comes next |
Bug reports, focused feature proposals and compatibility fixtures are welcome. Please run:
npm ci --ignore-scripts
npm run checkbefore opening a PR. See CONTRIBUTING.md.
The current published release is shown by the GitHub and npm badges above. The v1.0.0 contract froze scenario version: 1, core run result schemaVersion: 1, the public package entry point and the documented CLI command names. Future incompatible schema changes must use an explicit new schema version and migration path rather than silently reinterpreting v1 data.
MIT. See LICENSE.
Claude and Claude Code are products/trademarks of Anthropic. Claude Code Canary is an independent open-source project and is not affiliated with or endorsed by Anthropic.