The reusable EvalShift GitHub Action runs the CLI in CI, pushes the run to hosted EvalShift, compares it with a compatible base-branch run, and updates the pull request with a regression summary.
From a new EvalShift project:
evalshift init --ci # capture-first projectThis writes .github/workflows/evalshift.yml — a production-shaped,
self-documenting workflow (its header comment carries the full setup
checklist) with three jobs:
discover— lists every committed suite under.evalshift/suites/*/golden.jsonl. A suite added byevalshift capture syncis evaluated on the next run with no workflow edit; a project with no suites yet skips green with a notice instead of failing.eval <suite>— a matrix job per discovered suite (the action evaluates one suite per invocation), selected withsuite-name:— the key the suite is wired under inevalshift.yaml, not its path, so the suite's own evaluator block travels with it (see Selecting a suite). Runsfail-on: policy, so the verdict is the one this run computed against its ownmigration_policyblock inevalshift.yamland pushed inside the bundle.evalshift-versionis pinned to the CLI that scaffolded the project: the CLI that reads the config in CI must be at least as new as the CLI that wrote it locally (extra: forbidrejects newer keys), and the CLI warns when the pin falls behind (or runs ahead of the local CLI) — see Pin drift.max-paralleldefaults to 1 — raise it toward your hosted plan's in-flight-run ceiling (Free 1, Pro 5, Team 10). The PR comment is posted by the first matrix job only: the comment marker is a constant, so multiple suites would overwrite one another's summary.evalshift gate— the join job to require in branch protection. It fails if any suite failed and passes when evaluation was skipped (fork PR — no secret access, no suites committed, orEVALSHIFT_TOKENnot set yet). Require this check, not the per-suite jobs (dynamic names) and not theevalshift/regressioncommit status (with several suites the last writer wins that status).
The workflow keys off ${{ secrets.<PROVIDER>_API_KEY }} for the provider
chosen at init time; add further keys under the eval job's env: if your
judge or embedding models live in another family. Suites must be committed
for CI to see them — keep runtime data ignored and un-ignore just the suites:
.evalshift/*
!.evalshift/suites/
!.evalshift/toolsets/Runs on pushes to the main branch create the base-branch baselines pull
requests diff against; the workflow's concurrency block therefore cancels
superseded runs on PRs only, never on main.
The scaffold refuses to overwrite existing files. If you already have
an evalshift.yaml, run evalshift init --ci --directory somewhere scratch
and copy .github/workflows/evalshift.yml across instead of overwriting
your project.
EVALSHIFT_TOKEN: a service account key from the web app (Settings → API tokens → Service accounts), scoped torun:create+run:read+policy:read. Not a personal token — that one dies with its owner's membership and takes the pipeline with it.policy:readis what lets the defaultfail-on: policygate read the hosted verdict; without it the check falls back tofail-on: regressionsilently. A scoped key cannot auto-create the hosted project (project:createis owner-only), so create the project once in the web app and setcreate-project: false.- Provider API keys used by the source, target, judge, or embedding models.
Do not hard-code tokens in workflow YAML. Keep them in GitHub encrypted secrets
— repository or, better for production repos, environment secrets. Never expose
them to a pull_request_target workflow: that trigger runs the base repo's
workflow with secrets in scope against fork code.
Rotate on a schedule: rotate the key in the web app (the old one keeps working for a 24-hour grace window), update the GitHub secret, confirm a green run, then let the old key expire.
The Action expects the CLI project to know where hosted runs should land:
project: acme/model-migrationThe default workflow runs the local pipeline, finds the latest run id, and pushes that run to hosted EvalShift.
The action passes the hosted token through environment variables so command output does not expose it.
On pull requests, the Action:
- Pushes the candidate run.
- Looks for the latest compatible run on the base branch.
- Fetches the hosted diff if a baseline exists.
- Creates or updates one PR comment marked by EvalShift.
- Sets commit status
evalshift/regression.
If no compatible baseline exists, the comment explains that the run was pushed
but there is no baseline yet. Under regression / any-slice-regression,
gating passes. Under the default policy mode, the job still follows the
run's policy verdict.
| Mode | Behavior |
|---|---|
policy (default) |
Ask hosted EvalShift for the migration-policy verdict — the verdict the run itself computed against the migration_policy limits in evalshift.yaml and carried in its bundle, returned rather than re-scored. fail fails; pass/conditional_pass/inconclusive pass. A run pushed with no migration_policy is reported as ungated with a workflow warning and passes, unless require-policy: true. If the policy check is unreachable, falls back to regression gating and says so. |
never |
Do not fail the workflow for hosted regressions. |
regression |
Fail when the hosted diff reports one or more regressed examples. |
any-slice-regression |
Fail when any slice pass rate moves down. |
The Action can still fail for setup errors, provider auth errors, invalid config, upload failures, or finalize failures.
Common inputs:
| Input | Default | Description |
|---|---|---|
token |
required | Hosted EvalShift API token. |
host |
hosted default | Hosted API base URL. |
config |
evalshift.yaml |
Config path. |
suite-name |
— | Name of a suite wired under suites: in evalshift.yaml. Preferred — see Selecting a suite. Needs a CLI pin of 0.14.0 or newer. |
suite |
golden.jsonl |
Suite path, for a file that is not wired into the config (one suite per invocation). Mutually exclusive with suite-name. |
fail-on |
policy |
Gate mode — see the table above. |
require-policy |
false |
Fail the job when the pushed run carries no migration policy (policy mode only). |
evalshift-version |
action default (may lag) | Exact CLI version installed from PyPI. Always set it: it must be at least as new as the CLI that writes your evalshift.yaml (reader ≥ writer). init --ci pins it to the scaffolding CLI. |
create-project |
true |
Allow project auto-create when permissions allow it. |
comment |
true |
Post or update the PR comment on pull requests. |
See the action repository README for the full input list.
The action takes either suite-name: (a key under suites: in
evalshift.yaml) or suite: (a path). They load the same rows, but only the
name resolves that suite's own evaluators: block — the one evalshift capture sync writes for a tool-calling suite:
suites:
planner:
source: captured
path: .evalshift/suites/planner/golden.jsonl
evaluators:
tool_selection:
- name: routing
conformance: expected
divergence: setSelect that suite by path and it is scored with the top-level evaluators:
instead. There is no warning — a bare path is a legitimate way to run a suite
that has no entry under suites:. If the top level is semantic + llm_judge
and the suite's rows are tool calls, nothing scores at all and the run fails at
analyze with scores.jsonl is empty.
So: a suite with an entry under suites: is selected by name, which is what
init --ci scaffolds. suite: is for a one-off file outside the config.
evalshift.yaml rejects unknown keys (extra: forbid), so a config written by a
newer CLI can fail outright on an older one. The rule is reader ≥ writer:
the version the action installs in CI must be at least as new as the CLI you
run capture sync and init with locally. Upgrading locally without bumping
evalshift-version is the common way to break this.
The CLI checks for it wherever it writes or validates config — capture sync,
init (without --ci, next to a workflow it did not write; init --ci pins
the scaffolding CLI itself and does not warn about the file it just wrote),
doctor (a ci pin row), and validate. It parses every .github/workflows/*.yml for
evalshift/evalshift-action steps and compares their evalshift-version
with its own:
⚠ CI installs evalshift 0.12.1 (.github/workflows/evalshift.yml, job evalshift) but the local CLI is 0.13.1 — an older CLI rejects config keys a newer one writes.
Fix: set `evalshift-version: "0.13.1"` on the evalshift/evalshift-action step.
| Status | Trigger | Fix |
|---|---|---|
| stale | a literal pin is older than the local CLI | set evalshift-version: "<local>" |
| unpinned | a step has no evalshift-version (action default applies and may lag) |
add the pin |
| ahead | every pin is newer than the local CLI | pip install -U evalshift |
Equal pins, ${{ }} expressions, unparseable values, and an editable install
without metadata are silent. The check is advisory only: the CLI never edits
your workflow and never changes an exit code, and in CI it is a no-op by
construction — the running CLI is the pin. Config version: 1 is not
bumped for additive fields; see the
config version policy.
| Symptom | Likely cause | Fix |
|---|---|---|
| No PR comment | Missing pull-requests: write or issues: write. |
Add both permissions to the workflow. |
| No commit status | Missing statuses: write. |
Add the permission. |
| Push fails with missing project | Project does not exist and token cannot auto-create it. | Create the project in the web app or use an org-scoped owner token for first setup. |
| No compatible baseline | Base branch has not pushed a compatible run yet. | Merge or run EvalShift on the base branch once. |