docs: correct stale facts found by a docs-currency audit - #23
Merged
Merged
Conversation
DOCS.md still advertised 1.0.1 after 1.1.0 shipped, while llms-full.txt had the right number. Both headers are now checked against evalshift_cli.__version__, the installed metadata of pyproject's version. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The docstring described filter as a Python expression evaluated against each example's inputs. It is a literal tag, and analysis does not read the top-level slices: block at all: slices come from example tags and per-slice budgets are keyed by tag under migration_policy.slices. Docstring only; no behaviour change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Audit findings F1-F26 against the code, mirrored across both references: the init profile table (INIT_PROFILE_POLICIES numbers plus a tool-divergence column), a multi-turn example that now validates, the top-level slices: block documented as validated but not applied, resume hashing the suite path rather than its contents, the response cache serving tool-less examples only (with the full key), compare opening the report only with --open, the CLI import name and single-environment install, the complete failure-label set, optional config version, --no-insights skipping silently, --policy-gate failing with no policy, EVALSHIFT_DIR's full reach, the action's full input list and require-policy, exit code 2, and login reusing a still-valid token. Twins from the pages audit are fixed here too: imported traces stay local, push reuses an existing bundle, bundle needs a git SHA, and failed calls become errored rows excluded from the statistics. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Audit findings P1-P28 plus the twins of the reference fixes: imported traces are not uploaded, the bundle's trace streams carry no tool results or model_call events, slices come from tags (the example configs' slices: blocks are annotated as not applied), the tool-less-only cache, resume and policy-gate semantics, action no-baseline and require-policy behaviour, record_model_call's required tools=, suite auto-selection, errored rows, ls -t for the newest run, init --provider, EVALSHIFT_NONINTERACTIVE, the git-SHA requirement, push reusing an existing bundle, and a housekeeping command list. CHANGELOG gains one Unreleased Fixed entry. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`pre-commit run --all-files` without --hook-stage runs only the commit-stage hooks, the same false "everything" claim already corrected in AGENTS.md. Also re-wraps README.md's never-uploads bullet and the migration_policy.slices paragraph in docs/configuration.md to the surrounding line width. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The top-level slices: block has no effect on analysis, but the bundle's eval_config_hash is computed over the evaluator_config snapshot that includes it, and hosted EvalShift only pairs a candidate with a baseline of equal eval_config_hash. Editing or removing the block therefore breaks baseline compatibility with earlier hosted runs; every "not applied" site now says so, and docs/hosted.md marks the recorded slice definitions as not applied. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
utils/ci_pin.py reports three statuses -- stale, unpinned and ahead (every pin newer than the local CLI) -- and capture sync, init, validate and doctor print whichever finding check_ci_pin returns. The Pin drift sections already said so; the per-command summaries still described older or missing pins only. They now cover newer pins too, with the fix the ahead finding prints (pip install -U evalshift), and the configuration.md paragraph keeps its reader >= writer argument while explaining why an ahead pin still warns. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Sep 30, 2026
babaliauskas
added a commit
that referenced
this pull request
Oct 1, 2026
Revert the "tool-less examples only / agent suites run live at full price" statements that PR #23 added across DOCS.md, llms-full.txt and docs/{configuration,evaluators,getting-started,faq}.md, and describe the per-round key instead. docs/agents.md gains a bullet on how replayed rounds cache. CHANGELOG: a Changed entry, and the unreleased docs-audit entry now says the tool-less limit was the old behaviour. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes 55 places where the CLI docs disagreed with the code. They're spread across DOCS.md, llms-full.txt, README, AGENTS.md, CONTRIBUTING.md,
docs/*.mdandexamples/.Two audits checked every claim against
src/and--helpoutput, and two claims against the server. A fixer re-verified each finding before editing, and a reviewer checked every fix. A few audit findings were corrected along the way. For example, the methodology page's "four budgets" is three on the server.CLI behaviour is unchanged. The only non-doc changes are the
SliceConfigdocstring and one new guard test.Worth knowing
slices:config block does nothing today. It's validated and recorded in the bundle, but analysis never reads it. Every example tag becomes its own slice automatically, and per-slice budgets are keyed by tag undermigration_policy.slices. The docs now say so. Editing the block still changeseval_config_hash, which breaks hosted baseline matching with earlier runs. Whether to wire it up or remove it is a separate decision.runof an agent suite is live and full price. The docs are now scoped to tool-less examples.traces importdata is uploaded with the bundle; it isn't. The bundle carries only the replay's own tool-call traces.Also corrected
Wrong values
--profiletable: the model-upgrade numbers were the unused preset. It also gains a tool-divergence column.TOOL_GROUND_TRUTH_MISS.Broken examples
TypeError:record_model_callneedstools=.ls .evalshift/runs | head -1picked the oldest run, not the newest.Commands and gating
--policy-gatealso fails when no policy is configured, andinconclusivepasses.--gateis now explained.compareonly opens the report with--open.--no-insightsskips silently.push <run-id>reuses an existing bundle, andbundlerequires git orGITHUB_SHA.EVALSHIFT_NONINTERACTIVEand the housekeeping commands.Action docs
require-policy.Setup and dev workflow
pre-commit run --all-filesruns only the commit-stage hooks;make ciis the CI mirror.EVALSHIFT_DIRcovers suites and toolsets as well as captures.Leftovers
all-alias references.Checks
make ciis green: ruff, format, mypy --strict, 2351 tests and 94.44% coverage. The CHANGELOG has one### Fixedbullet under[Unreleased]. Nothing underdocs/superpowers/or in past CHANGELOG sections changed.🤖 Generated with Claude Code