Skip to content

docs: correct stale facts found by a docs-currency audit - #23

Merged
babaliauskas merged 7 commits into
mainfrom
docs/cli-docs-currency
Sep 30, 2026
Merged

babaliauskas merged 7 commits into
mainfrom
docs/cli-docs-currency

Conversation

@babaliauskas

Copy link
Copy Markdown
Collaborator

Fixes 55 places where the CLI docs disagreed with the code. They're spread across DOCS.md, llms-full.txt, README, AGENTS.md, CONTRIBUTING.md, docs/*.md and examples/.

Two audits checked every claim against src/ and --help output, and two claims against the server. A fixer re-verified each finding before editing, and a reviewer checked every fix. A few audit findings were corrected along the way. For example, the methodology page's "four budgets" is three on the server.

CLI behaviour is unchanged. The only non-doc changes are the SliceConfig docstring and one new guard test.

Worth knowing

  • The slices: config block does nothing today. It's validated and recorded in the bundle, but analysis never reads it. Every example tag becomes its own slice automatically, and per-slice budgets are keyed by tag under migration_policy.slices. The docs now say so. Editing the block still changes eval_config_hash, which breaks hosted baseline matching with earlier runs. Whether to wire it up or remove it is a separate decision.
  • Tool-calling examples aren't cached. The docs promised free repeat runs, but every run of an agent suite is live and full price. The docs are now scoped to tool-less examples.
  • Imported agent traces stay local. Three pages said traces import data is uploaded with the bundle; it isn't. The bundle carries only the replay's own tool-call traces.

Also corrected

Wrong values

  • The --profile table: the model-upgrade numbers were the unused preset. It also gains a tool-divergence column.
  • DOCS.md's version: 1.0.1 → 1.1.0. A new test now pins it to the package version.
  • The "complete" failure-label list was missing TOOL_GROUND_TRUTH_MISS.

Broken examples

  • The multi-turn golden example didn't validate.
  • Two SDK examples raised TypeError: record_model_call needs tools=.
  • ls .evalshift/runs | head -1 picked the oldest run, not the newest.

Commands and gating

  • Resume doesn't hash suite contents, so an edited suite at the same path resumes silently.
  • --policy-gate also fails when no policy is configured, and inconclusive passes. --gate is now explained.
  • compare only opens the report with --open.
  • --no-insights skips silently.
  • Exit code 2 is now documented.
  • push <run-id> reuses an existing bundle, and bundle requires git or GITHUB_SHA.
  • Suite auto-select is documented, along with EVALSHIFT_NONINTERACTIVE and the housekeeping commands.

Action docs

  • Now cover all inputs, including require-policy.
  • Policy-mode gating with no baseline is corrected.
  • The CI-pin warnings now include pins that are newer than the local CLI.

Setup and dev workflow

  • CLI and SDK install into one environment, not two.
  • pre-commit run --all-files runs only the commit-stage hooks; make ci is the CI mirror.
  • EVALSHIFT_DIR covers suites and toolsets as well as captures.

Leftovers

  • The last all-alias references.

Checks

make ci is green: ruff, format, mypy --strict, 2351 tests and 94.44% coverage. The CHANGELOG has one ### Fixed bullet under [Unreleased]. Nothing under docs/superpowers/ or in past CHANGELOG sections changed.

🤖 Generated with Claude Code

babaliauskas and others added 7 commits September 30, 2026 12:46
DOCS.md still advertised 1.0.1 after 1.1.0 shipped, while llms-full.txt
had the right number. Both headers are now checked against
evalshift_cli.__version__, the installed metadata of pyproject's version.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The docstring described filter as a Python expression evaluated against
each example's inputs. It is a literal tag, and analysis does not read the
top-level slices: block at all: slices come from example tags and per-slice
budgets are keyed by tag under migration_policy.slices. Docstring only; no
behaviour change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Audit findings F1-F26 against the code, mirrored across both references:
the init profile table (INIT_PROFILE_POLICIES numbers plus a
tool-divergence column), a multi-turn example that now validates, the
top-level slices: block documented as validated but not applied, resume
hashing the suite path rather than its contents, the response cache
serving tool-less examples only (with the full key), compare opening the
report only with --open, the CLI import name and single-environment
install, the complete failure-label set, optional config version,
--no-insights skipping silently, --policy-gate failing with no policy,
EVALSHIFT_DIR's full reach, the action's full input list and
require-policy, exit code 2, and login reusing a still-valid token.
Twins from the pages audit are fixed here too: imported traces stay
local, push reuses an existing bundle, bundle needs a git SHA, and
failed calls become errored rows excluded from the statistics.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Audit findings P1-P28 plus the twins of the reference fixes: imported
traces are not uploaded, the bundle's trace streams carry no tool results
or model_call events, slices come from tags (the example configs' slices:
blocks are annotated as not applied), the tool-less-only cache, resume and
policy-gate semantics, action no-baseline and require-policy behaviour,
record_model_call's required tools=, suite auto-selection, errored rows,
ls -t for the newest run, init --provider, EVALSHIFT_NONINTERACTIVE, the
git-SHA requirement, push reusing an existing bundle, and a housekeeping
command list. CHANGELOG gains one Unreleased Fixed entry.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`pre-commit run --all-files` without --hook-stage runs only the
commit-stage hooks, the same false "everything" claim already corrected
in AGENTS.md. Also re-wraps README.md's never-uploads bullet and the
migration_policy.slices paragraph in docs/configuration.md to the
surrounding line width.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The top-level slices: block has no effect on analysis, but the bundle's
eval_config_hash is computed over the evaluator_config snapshot that
includes it, and hosted EvalShift only pairs a candidate with a baseline
of equal eval_config_hash. Editing or removing the block therefore breaks
baseline compatibility with earlier hosted runs; every "not applied"
site now says so, and docs/hosted.md marks the recorded slice
definitions as not applied.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
utils/ci_pin.py reports three statuses -- stale, unpinned and ahead
(every pin newer than the local CLI) -- and capture sync, init, validate
and doctor print whichever finding check_ci_pin returns. The Pin drift
sections already said so; the per-command summaries still described
older or missing pins only. They now cover newer pins too, with the fix
the ahead finding prints (pip install -U evalshift), and the
configuration.md paragraph keeps its reader >= writer argument while
explaining why an ahead pin still warns.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@babaliauskas
babaliauskas merged commit af4e6fe into main Sep 30, 2026
4 checks passed
babaliauskas added a commit that referenced this pull request Oct 1, 2026
Revert the "tool-less examples only / agent suites run live at full
price" statements that PR #23 added across DOCS.md, llms-full.txt and
docs/{configuration,evaluators,getting-started,faq}.md, and describe the
per-round key instead. docs/agents.md gains a bullet on how replayed
rounds cache. CHANGELOG: a Changed entry, and the unreleased docs-audit
entry now says the tool-less limit was the old behaviour.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant