Status boardFree
Consumer-product ratings on the six-dimension rubric (rebuilt July 2026). See the full model ratingsfor each lab’s per-dimension breakdown.
Policy changes · summary, severity & whyPro
When a lab changes a policy, a plain-language summary with a severity rating and the reasoning — across all six labs.
Unlock the full analysis
Everything in the free overview, plus what members get:
- ✓Every relevant frontier paper, rated — evidence quality, significance & SI domain
- ✓A plain-language note on each paper
- ✓Policy-change analysis across all six labs — summary, severity & the reasoning
- ✓Email alerts when a lab changes its terms — coming soon
- ✓Word-level redline diffs — coming soon
- ✓Searchable history & archive — coming soon
$70/ year
≈ $5.83/mo · 2 months free
No account needed to subscribe — we’ll set one up from your email. Cancel anytime.
Relevant papersFree
BayesAME: Bayesian Active Model EvaluationarXiv · 1d ago
Lottery Tickets Are Not Deployment TicketsarXiv · 1d ago
Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous DataarXiv · 1d ago
PIKS: Universal Physics-Informed Kernel MethodsarXiv · 1d ago
Visual Credit Audit for Multimodal Spatial ReasoningarXiv · 1d ago
Tracked benchmarksFree
IG-Bench (IdeaGene-Bench)27.3%Scientific lineage reasoning and lineage-grounded idea generation — can AI trace how research ideas evolve and propose a coherent next step. A direct probe of AI-conducted science.Recursive Improvement · Capability Trajectory · best 27.3% (2026-07-09) · unreplicated — evidence, not an input
METR Task-Completion Time Horizons17hHow long a real software/research task (in human-expert hours) a frontier agent can complete autonomously at 50% reliability.Autonomy · Capability Trajectory · best 17h (2026-05-08) · unreplicated — evidence, not an input
ARC-AGI-37.8%Whether an agent can explore novel interactive environments, acquire goals, build world models, and adapt over long horizons without instructions.Autonomy · Reasoning Frontier · best 7.8% (2026-07-09) · unreplicated — evidence, not an input
SWE-bench Pro (Public)61.5%Resolving realistic long-horizon software-engineering issues (multi-file patches) in contamination-resistant repos with standardized scaffolding.Recursive Improvement · Capability Trajectory · best 61.5% (2026-07-12) · unreplicated — evidence, not an input
Each paper rated · evidence, significance, SI domain & a notePro
Every relevant paper rated for evidence quality, significance, and SI domain — each with a plain-language note.
Unlock the full analysis
Everything in the free overview, plus what members get:
- ✓Every relevant frontier paper, rated — evidence quality, significance & SI domain
- ✓A plain-language note on each paper
- ✓Policy-change analysis across all six labs — summary, severity & the reasoning
- ✓Email alerts when a lab changes its terms — coming soon
- ✓Word-level redline diffs — coming soon
- ✓Searchable history & archive — coming soon
$70/ year
≈ $5.83/mo · 2 months free
No account needed to subscribe — we’ll set one up from your email. Cancel anytime.