What a benchmark scores with nothing in it: content-free members against each task's own majority baseline, plus a judged arm and a judge-disagreement control. Preregistered. WIP.
-
Updated
Sep 17, 2026 - Python
What a benchmark scores with nothing in it: content-free members against each task's own majority baseline, plus a judged arm and a judge-disagreement control. Preregistered. WIP.
Reproducible code and metrics for measuring exploitability and contamination in LLM evaluation harnesses.
To associate your repository with the evaluation-integrity topic, visit your repo's landing page and select "manage topics."