Skip to content
#

unsupervised-evaluation

Here are 2 public repositories matching this topic...

Research testbed pairing allotment (maximin stratified sortition, Flanigan et al. Nature 2021) with ntqr algebraic logic for ground-truth-free evaluation of noisy binary classifiers — a 96-seed grid over 3/6/9/12-judge panels and 300-item corpora, validated against live gemma3:4b reviewers via Ollama.

  • Updated Aug 29, 2026
  • Python

Label-free evaluation of LLM judge ensembles: 3 market personas across 6 local Ollama models vote on 64 scenarios, and the NTQR 0.8 algebraic-geometry solver recovers per-persona accuracy from votes alone — no answer key. Includes error-independence controls, bootstrap 95% CIs, and results-injected manuscript builds.

  • Updated Aug 29, 2026
  • Python

Add this topic to your repo

To associate your repository with the unsupervised-evaluation topic, visit your repo's landing page and select "manage topics."

Learn more