A reporting layer
over evaluation
infrastructure.
Evaluation Cards is a collection of reported model–benchmark results, organized under a five-level rollout hierarchy and four interpretive signals computed over the joined record.
Interpretive signals
Four signals computed over each (model, benchmark, metric-path) record and aggregated to the corpus level. Per-record instances appear on every model and benchmark page.
of reported scores have a complete setup recorded. The rest cannot be independently re-run.
92% have at least one undocumented field. Most often missing: temperature (90%), max tokens (89%).
mean across 86,304 reported score triples.
Observed range: 7% to 93%.
of reported score triples have reports from more than one party.
89% third-party, 7% first-party of 86,304 unique triples.
of setup-eligible groups diverge across variants (531 of 2,638).
Cross-party divergence: 22%.
Benchmark families
All 87 →Mercor ACE
1 reported benchmark across this family.
AgentHarm
1 reported benchmark across this family.
How Inference Compute Shapes Frontier LLM Evaluation
7 reported benchmarks across this family.
Alpaca-EVAL-V1
1 reported benchmark across this family.
Alpaca-EVAL-V2
1 reported benchmark across this family.
AlpacaEval
4 reported benchmarks across this family.
Every score resolves to an explicit path through this hierarchy, so aggregate claims drill down to the evidence supporting them.