Applied mathematician (PhD). I build systems that are honest about what they don't know.
I work on LLM evaluation and document AI: evaluation harnesses with real statistics, golden datasets with documented rubrics, and extraction pipelines whose accuracy is measured, not asserted. Four years shipping production models at a Fortune-100 insurer; most recently a document-AI research contract. Currently available for evaluation and document-AI work.
- eval-toolkit — evaluation library on PyPI: bootstrap confidence intervals, leakage checks, versioned result schemas, CI gates.
- prompt-injection-detection-prototype — a full detection study published with its negative result: confidence intervals, baseline contamination disclosed, failure modes analyzed. Honest evaluation, demonstrated.
- ir-eval — statistical retrieval evaluation for CI/CD: paired tests and drift detection, because most RAG systems fail silently.
- research-kb — a multi-thousand-source research knowledge base: PDF ingestion → hybrid retrieval (BM25 + vectors + reranking) → MCP server, with a retrieval eval suite gating the weekly rebuild.
- temporalcv — released Python package for time-series cross-validation with gap enforcement (the leakage everyone ships).
Site: brandon-behring.dev



