4 papers
Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
Zhiqin Yang, Jingwen Fu, Yuhan Liu +16
Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code,…
Pose-Agnostic Robotic Functional Grasping via Observation-Action Canonicalization
Le Qiu, Cole Harrison, Jiankai Sun +5
Functional robotic grasping requires a policy that generalizes across diverse object geometries and poses while maintaining task-specific contact precision. We study this challenge…
SARM2: Multi-Task Stage Aware Reward Modeling for Self Improving Robotic Manipulation
Qianzhong Chen, Hau Zheng, Justin Yu +8
Fine-tuning vision-language-action (VLA) policies for long-horizon manipulation still relies heavily on behavior cloning, which requires costly high-quality demonstrations and keep…
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
Rishi Desai, Jesse Hu, Joan Cabezas +23
AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments. Yet current agent b…