Certified reference models for AI evals: deterministically broken in one documented way each, so a benchmark's detection rate is measurable against ground truth.
-
Updated
Aug 28, 2026 - Python
Certified reference models for AI evals: deterministically broken in one documented way each, so a benchmark's detection rate is measurable against ground truth.
To associate your repository with the defect-injection topic, visit your repo's landing page and select "manage topics."