A reporting layer
over evaluation
infrastructure.

Evaluation Cards is a collection of reported model–benchmark results, organized under a five-level rollout hierarchy and four interpretive signals computed over the joined record.

Corpus snapshot · September 7, 2026
8,834
Models
Tracked across reporting sources
501,917
Reported results
(model, benchmark, metric) triples
48
Reporting organizations
Distinct evaluator initiatives in this corpus
1,637
Model developers
Distinct model-publishing organizations
86
Benchmark families
Top of the rollout hierarchy
1,396
Single benchmarks
1,831 slices · 1,929 metrics

Interpretive signals

Four signals computed over each (model, benchmark, metric-path) record and aggregated to the corpus level. Per-record instances appear on every model and benchmark page.

Benchmark families

All 86
Five-level rollout hierarchy

Every score resolves to an explicit path through this hierarchy, so aggregate claims drill down to the evidence supporting them.