Decision-safe evaluation + Streamlit dashboard for AI vs Human vs Post-Edited AI text detection. Generates a reliability report card (Accuracy, Macro F1, ECE, Brier), calibration plots, confidence histograms, and a coverage-vs-performance abstention curve. Recommends an operating threshold for human-review routing.
nlp coverage machine-learning text-classification reliability plotly calibration confusion-matrix thresholding uncertainty-estimation model-evaluation responsible-ai streamlit abstention selective-classification content-authenticity ai-detection brier-score expected-calibration-error decision-safety
-
Updated
Sep 12, 2026 - Python