2 papers
cs.AI2026
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
Yiwei Li, Wanli Yang, Hexiang Tan +10
Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this…
cs.AI2026
Offline-Online Curriculum RL for Multimodal Reasoning
Wendi Deng, Hang Du, Guoshun Nan +11
Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correct final answers. This behavior undermines…