Pinned Loading
-
llm-evaluation-framework
llm-evaluation-framework PublicProduction-grade evaluation framework for LLMs, RAG pipelines, and autonomous AI agents with deterministic checks, LLM-as-a-Judge, RAG Triad, and an interactive observatory dashboard.
Python
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.