Capability Watch

Papers

Every frontier-AI research paper we’ve surfaced from arXiv, newest first — the raw capability signals Frontier Watch tracks. Scroll to load older ones.

880 papers tracked · ← dashboard

  1. 1BayesAME: Bayesian Active Model Evaluation
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  2. 2TreeCCA: Canonical Correlation Analysis via Gradient-Boosted Trees
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  3. 3HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier Models
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  4. 4Lottery Tickets Are Not Deployment Tickets
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  5. 5Mitigating Compounding Error via Video Representation Regularization
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  6. 6Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  7. 7PIKS: Universal Physics-Informed Kernel Methods
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  8. 8Visual Credit Audit for Multimodal Spatial Reasoning
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  9. 9Equilibrium Training of Energy-Based Models with Parallel Trajectory Tempering
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  10. 10MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  11. 11On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  12. 12Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  13. 13InferScale: GPU-Native KV Injection for Personalized LLM Serving
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  14. 14Sky sphere representation in language models
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  15. 15AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  16. 16Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  17. 17Linguistic Monoculture in LLM-Assisted Language Use
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  18. 18DLAM: Distributional Latent Actions with Temporal Constraints
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  19. 19Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  20. 20MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  21. 21OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  22. 22SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  23. 23When Do Learned Diffusion Proposals Help Constraint Solving? A Controlled Study on Continuous Algebraic Systems
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  24. 24Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  25. 25APEX-Accounting
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  26. 26Can AI agents conduct open-ended AI research? Early evidence from two case studies
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  27. 27Mental World Modeling
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  28. 28Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-29
  29. 29Stemma: Induced Decision Regions Reveal LLM Provenance
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
  30. 30AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
  31. 31RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
  32. 32Distributing Security Controls Through Harness Engineering
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
  33. 33Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
  34. 34HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
  35. 35Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
  36. 36Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
  37. 37SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
  38. 38Penelope: Localized Latent Recurrence for Efficient Structured Reasoning
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
  39. 39dtControl2+$\varepsilon$: Trading Optimality for Explainability in MDPs via Decision Trees
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28
  40. 40Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
    arXiv cs.AI / cs.LG / cs.CL (recent) · 2026-07-28