Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

AMBER — Archived-Moment Behavioral Evaluation Replay

封进琥珀,重做当时的题。 — pack the moment in amber; work the problem of that time again.

AMBER is a formal evaluation method: take a real, auditable historical event; restore its exact pre-outcome state; give the candidate only cutoff-available information; physically seal all later evidence and evaluator materials; let the candidate diagnose, decide, act, abstain, refuse, or escalate; judge against criteria fixed before any output is viewed; preserve everything for audit.

Why: public benchmarks are static public question sets — training contamination is common and unauditable, scores inflate, and static Q&A does not measure what real work demands (multi-step diagnosis, abstention, refusal, escalation under uncertainty). AMBER replays real events whose leak status is checkable because the source is controlled. It runs alongside public benchmarks, never as their replacement.

Status

Draft v0.2.2. The normative Core is stable; the protocols/, schemas/, and profiles/ layers are inherited from the predecessor corpus and are not yet published here.

Contents

  • AMBER-Core-Specification.md — the normative cross-domain core: purpose, definition, mechanism, 8 invariants, 8 boundaries, epistemic limits, naming review, adoption rules.

License

Apache-2.0 — see LICENSE.

About

AMBER — Archived-Moment Behavioral Evaluation Replay: replay sealed real events to evaluate model/agent behavior. Core specification (draft v0.2.2).

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors