2 papers
cs.LG2026
R^3: Replay, Reflection, and Ranking Rewards for LLM Reinforcement Learning
Zhizheng Jiang, Kang Zhao, Weikai Xu +5
Large reasoning models (LRMs) aim to solve diverse and complex problems through structured reasoning. Recent advances in group-based policy optimization methods have shown promise…
cs.CV2025
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
Zixin Zhang, Kanghao Chen, Xingwang Lin +6
The ability to use, understand, and create tools is a hallmark of human intelligence, enabling sophisticated interaction with the physical world. For any general-purpose intelligen…