An End-to-End Evaluation Framework for Entity Resolution Systems
-
Updated
Dec 3, 2023 - Python
An End-to-End Evaluation Framework for Entity Resolution Systems
Benchmark your MCP server.
TorchMetrics-based evaluation metrics for robotics policies and robot learning
A single-file Python CLI that pre-registers AI/ML accuracy claims with SHA-256. Lock the threshold before the data, or it didn't happen.
Live index of LLM evaluation tools and benchmarks, refreshed every 15 minutes from GitHub
✍️ Collaborate on writing technical content for the Giskard Community
Measurement-disciplined optimizer for Claude Agent Skills — three-gate system (stability, effect size, function preservation), anchored rubric, train/holdout split. Inspired by Karpathy's autoresearch and alchaincyf's darwin-skill.
An open-source Streamlit web app to generate beautiful confusion matrices for multi-class machine learning models. Supports numeric and string labels, CSV upload, manual label entry, custom color maps, and displays evaluation metrics like Accuracy, Precision, Recall, and F1-score. Users can download the confusion matrix as an image.
Safety-first legal NLP system with hierarchical long-document processing, deterministic inference, clause extraction, and rule-based risk engine — built for traceability and deployment constraints.
GitHub Action for SWE-bench Pro evaluation powered by mcpbr
Enterprise-grade machine learning observability platform that detects data drift, concept drift, and performance degradation in production models. Features statistical drift detection (KS test, PSI), real-time alerting, Redis caching, and FastAPI backend.
Data Science Challenge from Coursera Project : Loan Default Prediction
This project contains codes and paperwork based on the course CSI5155 at University of Ottawa (delivered by Professor Dr. Herna Viktor).
Public scorecard of how 27 ML eval claims meet 9 PRML falsifiability criteria. CC0 data; MIT tooling.
A decision-oriented benchmark framework for evaluating action-conditioned world models beyond static AI benchmarks.
AI quality dashboard for two production scoring systems. PostgreSQL, FastAPI, React. Finds and quantifies five systematic scoring failures.
Config-driven tabular ML benchmark comparing preprocessing, feature engineering, tuning, and AutoML-style baselines under shift.
Self-hosted accuracy evals + correction loop for document extraction. Bring your own model.
Field manual for the PRML v0.1 specification. 13 patterns, 4 anti-patterns, 4 working examples. CC0.
A MATLAB-based machine learning project that implements a Naive Bayes spam email classifier using the UCI Spambase dataset. Includes feature selection, model tuning, performance evaluation, and deployment-ready model export.
Add a description, image, and links to the ml-evaluation topic page so that developers can more easily learn about it.
To associate your repository with the ml-evaluation topic, visit your repo's landing page and select "manage topics."