Capability Watch
Papers
Every frontier-AI research paper we’ve surfaced from arXiv, newest first — the raw capability signals Frontier Watch tracks. Scroll to load older ones.
880 papers tracked · ← dashboard
- 1BayesAME: Bayesian Active Model EvaluationarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 2TreeCCA: Canonical Correlation Analysis via Gradient-Boosted TreesarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 3HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier ModelsarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 4Lottery Tickets Are Not Deployment TicketsarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 5Mitigating Compounding Error via Video Representation RegularizationarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 6Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous DataarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 7PIKS: Universal Physics-Informed Kernel MethodsarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 8Visual Credit Audit for Multimodal Spatial ReasoningarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 9Equilibrium Training of Energy-Based Models with Parallel Trajectory TemperingarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 10MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and RepairarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 11On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust RealignmentarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 12Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM AgentsarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 13InferScale: GPU-Native KV Injection for Personalized LLM ServingarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 14Sky sphere representation in language modelsarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 15AgentMap: Joint Equivalence and Subsumption Discovery for Ontology MatchingarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 16Minimal Markovization via Stable Quotients in Holonomy-Cover Decision ProcessesarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 17Linguistic Monoculture in LLM-Assisted Language UsearXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 18DLAM: Distributional Latent Actions with Temporal ConstraintsarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 19Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain BenchmarkarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 20MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program SynthesisarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 21OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic GroundingarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 22SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from ScratcharXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 23When Do Learned Diffusion Proposals Help Constraint Solving? A Controlled Study on Continuous Algebraic SystemsarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 24Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc TeamworkarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 25APEX-AccountingarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 26Can AI agents conduct open-ended AI research? Early evidence from two case studiesarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 27Mental World ModelingarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 28Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
- 29Stemma: Induced Decision Regions Reveal LLM ProvenancearXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
- 30AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal EvaluationarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
- 31RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-ImprovementarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
- 32Distributing Security Controls Through Harness EngineeringarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
- 33Messier: A High-Resolution Corpus for Cross-Benchmark Agent EvaluationarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
- 34HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data AlonearXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
- 35Interactive Reward Agent: GUI Task Evaluation via Environment-State VerificationarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
- 36Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language ModelsarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
- 37SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action ModelsarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
- 38Penelope: Localized Latent Recurrence for Efficient Structured ReasoningarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
- 39dtControl2+$\varepsilon$: Trading Optimality for Explainability in MDPs via Decision TreesarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
- 40Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical CasesarXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28