AIWG training-complete framework — corpus-to-dataset pipeline with SKILL.md agentic surface and optional Python runtime backend. Marketplace plugin for AIWG.
-
Updated
Apr 16, 2026 - Python
AIWG training-complete framework — corpus-to-dataset pipeline with SKILL.md agentic surface and optional Python runtime backend. Marketplace plugin for AIWG.
The official implementation repository for our EMNLP 2024 Findings paper, PaCoST: Paired Confidence Significance Testing for Benchmark Contamination Detection in Large Language Models.
Reproducible code and metrics for measuring exploitability and contamination in LLM evaluation harnesses.
Benchmark contamination and clustered inference in rare-disease gene prioritisation — a 1,047-case benchmark with per-item leakage flags, model-free floors, and fully re-derivable results.
Do quality filters pull benchmark questions into pretraining corpora? Injects MMLU, GSM8K, GPQA, ARC, HellaSwag, PIQA and TruthfulQA items into a corpus and ranks them with six quality classifiers, including DCLM, FineWeb-Edu and Gaperon.
API-only cross-lingual contamination detection tooling for TR-MMLU (Turkish MMLU benchmark for LLM evaluation)
Reproducible audit of public-to-renewal LLM benchmark score non-transfer
Erk-32B — Türkçe dil modeli (Qwen3-32B tabanlı): ölçüm protokolü, yeniden üretim betikleri, sonuç vektörleri. Ağırlıklar: huggingface.co/ecloudtech/Erk-32B
TurkishMMLU kirlilik listesi — Türkçe web'de birebir geçen 113 test sorusu (n=13), yöntem ve betik
Multi-benchmark contamination audit of Turkish LLM benchmarks (M1/M2/M3 probes, black-box API)
Detecting benchmark contamination by probing a language model's residual stream, with the controls that make the result mean something. Paper: arXiv:2608.12652
End-to-end Python research pipeline replicating "Beyond Benchmark Rankings": treats LLM selection as portfolio construction rather than leaderboard ranking. Estimates marginal utility, informational novelty, and correlated-failure risk per model, then runs budget-feasible greedy optimization to select LLM ensembles.
To associate your repository with the benchmark-contamination topic, visit your repo's landing page and select "manage topics."