AI-detector auditing toolkit: measures how often a detector flags genuine human writing, how stable its verdict is across seeds, and whether that verdict survives meaning-preserving edits. Detector-in-the-loop measurement harness. Claude Code skill + Python CLI. MIT.
python nlp open-source benchmark reproducible-research model-evaluation turnitin false-positives gptzero ai-detection zerogpt writing-tools ai-detector claude-code robustness-testing claude-code-skill unicode-security detector-evaluation
-
Updated
Sep 2, 2026 - Python