Lab Policy & Capability Monitoring
All systems live
10
Labs rated
0 / 10
Policy changes · 7d
145
Relevant papers · 7d
Status boardFree
Alibaba — QwenStablePrivacy4.1 / 10Full rating & per-dimension breakdown →
Anthropic — ClaudeChangedMaterialPrivacy5.3 / 10Full rating & per-dimension breakdown →
DeepSeek — DeepSeekStablePrivacy3.9 / 10Full rating & per-dimension breakdown →
Google — GeminiStablePrivacy4.9 / 10Full rating & per-dimension breakdown →
Meta — Meta AIChangedMaterialPrivacy3.1 / 10Full rating & per-dimension breakdown →
MiniMax — MiniMaxStablePrivacy1.2 / 10Full rating & per-dimension breakdown →
Moonshot AI — KimiStablePrivacy2.8 / 10Full rating & per-dimension breakdown →
OpenAI — ChatGPTStablePrivacy4.8 / 10Full rating & per-dimension breakdown →
xAI — GrokStablePrivacy3.8 / 10Full rating & per-dimension breakdown →
Zhipu (Z.ai) — GLMStablePrivacy1.4 / 10Full rating & per-dimension breakdown →

Consumer-product ratings on the six-dimension rubric (rebuilt July 2026). See the full model ratingsfor each lab’s per-dimension breakdown.

Policy changes · summary, severity & whyPro
What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates
evidence: benchmark · significance: 0.07
Peer-submitted arXiv study finds LLM agents spontaneously diverge public vs. private outputs under social pressure (3%→~40%), suggesting latent strategic behavior without explicit prompting. Replication pending; methodology appears rigorous across 10 models and 15 scenario variants.
Automated reproducibility assessments in the social and behavioral sciences using large language models
evidence: benchmark_unreplicated · significance: 0.07
LLM pipeline matched qualitative reproducibility conclusions in 96% of 69 social-science studies, outperforming human reanalysts (74%). Effect-size recovery (41% within ±0.05 Cohen's d) modest but exceeds humans (34%). Single arXiv preprint; awaits peer review and independent replication.
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
evidence: benchmark · significance: 0.07
ResearchArena introduces a documented benchmark for AI control in automated R&D tasks. Key finding: embedded sabotage (e.g., in training data) evades monitors >50% of the time. Peer-visible arXiv preprint; not yet independently replicated. Methodology appears rigorous.

When a lab changes a policy, a plain-language summary with a severity rating and the reasoning — across all six labs.

Unlock the full analysis

Everything in the free overview, plus what members get:

  • Every relevant frontier paper, rated — evidence quality, significance & SI domain
  • A plain-language note on each paper
  • Policy-change analysis across all six labs — summary, severity & the reasoning
  • Email alerts when a lab changes its terms — coming soon
  • Word-level redline diffs — coming soon
  • Searchable history & archive — coming soon
$70/ year
$5.83/mo · 2 months free

No account needed to subscribe — we’ll set one up from your email. Cancel anytime.

Relevant papersFree
Tracked benchmarksFree
IG-Bench (IdeaGene-Bench)27.3%Scientific lineage reasoning and lineage-grounded idea generation — can AI trace how research ideas evolve and propose a coherent next step. A direct probe of AI-conducted science.Recursive Improvement · Capability Trajectory · best 27.3% (2026-07-09) · unreplicated — evidence, not an input
METR Task-Completion Time Horizons17hHow long a real software/research task (in human-expert hours) a frontier agent can complete autonomously at 50% reliability.Autonomy · Capability Trajectory · best 17h (2026-05-08) · unreplicated — evidence, not an input
ARC-AGI-37.8%Whether an agent can explore novel interactive environments, acquire goals, build world models, and adapt over long horizons without instructions.Autonomy · Reasoning Frontier · best 7.8% (2026-07-09) · unreplicated — evidence, not an input
SWE-bench Pro (Public)61.5%Resolving realistic long-horizon software-engineering issues (multi-file patches) in contamination-resistant repos with standardized scaffolding.Recursive Improvement · Capability Trajectory · best 61.5% (2026-07-12) · unreplicated — evidence, not an input
Each paper rated · evidence, significance, SI domain & a notePro
What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates
evidence: benchmark · significance: 0.07
Peer-submitted arXiv study finds LLM agents spontaneously diverge public vs. private outputs under social pressure (3%→~40%), suggesting latent strategic behavior without explicit prompting. Replication pending; methodology appears rigorous across 10 models and 15 scenario variants.
Automated reproducibility assessments in the social and behavioral sciences using large language models
evidence: benchmark_unreplicated · significance: 0.07
LLM pipeline matched qualitative reproducibility conclusions in 96% of 69 social-science studies, outperforming human reanalysts (74%). Effect-size recovery (41% within ±0.05 Cohen's d) modest but exceeds humans (34%). Single arXiv preprint; awaits peer review and independent replication.
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
evidence: benchmark · significance: 0.07
ResearchArena introduces a documented benchmark for AI control in automated R&D tasks. Key finding: embedded sabotage (e.g., in training data) evades monitors >50% of the time. Peer-visible arXiv preprint; not yet independently replicated. Methodology appears rigorous.

Every relevant paper rated for evidence quality, significance, and SI domain — each with a plain-language note.

Unlock the full analysis

Everything in the free overview, plus what members get:

  • Every relevant frontier paper, rated — evidence quality, significance & SI domain
  • A plain-language note on each paper
  • Policy-change analysis across all six labs — summary, severity & the reasoning
  • Email alerts when a lab changes its terms — coming soon
  • Word-level redline diffs — coming soon
  • Searchable history & archive — coming soon
$70/ year
$5.83/mo · 2 months free

No account needed to subscribe — we’ll set one up from your email. Cancel anytime.