test: a committed mutant list, run nightly - #30
Merged
Merged
Conversation
My own mutation checks scored 14/14 on this package while an independent 54-mutant sweep found 27 survivors, including a checkpoint that recorded the newest row instead of the page boundary -- the defect 4.0.0 exists to fix -- and a "drift gate" that still passed with 30 of 36 columns dropped. A mutant list written by the author of the tests contains the mutations those tests already catch. That is grading my own homework, and the score was the reassuring kind. So the list is committed and reviewable: what is being checked, and more usefully what is not. Each entry carries why it matters, most of them naming a defect that actually reached a customer. Add one whenever a defect reaches main -- the mutation is the proof the new test would have caught it. Nightly and non-gating: it re-runs the suite once per mutant, which is minutes, and a survivor is information rather than a reason to block a merge. A mutant whose `find` no longer matches counts as a failure too, because a stale mutant has been silently testing nothing. Its first real run found two gaps, and this commit closes both. A budget exhausted on a RESUMED run deleted the archive. The `written == 0` deletion guard existed on the generic failure path and the budget path had the same test with nothing behind it -- and the budget path is the more dangerous of the two, because exit 75 is the ORDINARY outcome of a long back fill: the file vanishes, the next run appends only the tail and records complete. The line was added hours earlier while closing an "an empty file is a lie" finding, which is the same mistake made twice in a day: fix the site in front of me, do not look for its twin. Nothing published ever had the unguarded version. And `str(EntityType.SHOW)` is "EntityType.SHOW", so a customer filtering a CSV on SHOW matches nothing and is told nothing. The JavaScript SDK pins this; Python never did, because every stub in the file hands the writer a plain string and so never exercises the unwrap. One mutant also had to be split: a single `find` string matched both deletion sites, so it tested whichever came first while appearing to cover two. 19/19 killed, 336 tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
My own mutation checks scored 14/14 on this package while an independent 54-mutant sweep found 27 survivors — including a checkpoint that recorded the newest row instead of the page boundary (the defect 4.0.0 exists to fix) and a "drift gate" that still passed with 30 of 36 columns dropped.
A mutant list written by the author of the tests contains the mutations those tests already catch. That is grading my own homework, and the score was the reassuring kind.
So the list is committed and reviewable:
tests/mutation/mutants.json, each entry carrying why it matters, most of them naming a defect that actually reached a customer. The value is as much in what a reader can see is missing as in the score.Nightly and non-gating — it re-runs the suite once per mutant, which is minutes, and a survivor is information rather than a reason to block a merge. A mutant whose
findno longer matches counts as a failure too: a stale mutant has been silently testing nothing, which is the same failure mode as a test that cannot fail.Its first real run found two gaps
A budget exhausted on a resumed run deleted the archive. The
written == 0deletion guard existed on the generic failure path; the budget path had the same test with nothing behind it — and it is the more dangerous of the two, because exit 75 is the ordinary outcome of a long back fill. The file vanishes, the next run appends only the tail and recordscomplete: true.That line was added hours earlier while closing an "an empty file is a lie" review finding, which is the same mistake made twice in one day: fix the site in front of me, don't look for its twin. Nothing published ever had the unguarded version — 3.3.0's budget handler deletes nothing.
str(EntityType.SHOW)is'EntityType.SHOW', so a customer filtering a CSV onSHOWmatches nothing and is told nothing. The JavaScript SDK pins this; Python never did, because every stub in the file hands the writer a plain string and so never exercises the unwrap.One mutant also had to be split: a single
findmatched both deletion sites, so it tested whichever came first while appearing to cover two.19/19 killed, 336 tests.
🤖 Generated with Claude Code