Ship evals before you ship features.
-
Updated
Jun 22, 2026 - Nunjucks
Ship evals before you ship features.
Security working agreements for AI coding agents: hardened AGENTS.md, prompt/tool-injection guardrails, dependency hygiene, Scorecard-ready OSS setup
Agent-CE is a containerized continuous evaluation (CE) platform for web browsing agents. It provides production-ready Docker images and CI/CD pipelines for running and evaluating multiple agent frameworks including Browser Use, Notte, Anthropic Computer Use, and OpenAI Computer Use.
Continuous Evaluation Infrastructure for Production AI
Enterprise human intelligence infrastructure for AI systems — task generation, governed review pipelines, model integration, and continuous evaluation.
Is My Model Drifting? A public LLM regression tracker: a frozen, hash-fingerprinted suite run daily against 16 models across 5 labs, scoring accuracy, latency, verbosity, reliability and refusal rate. Deterministic grading, no LLM judge.
Add a description, image, and links to the continuous-evaluation topic page so that developers can more easily learn about it.
To associate your repository with the continuous-evaluation topic, visit your repo's landing page and select "manage topics."