7 papers
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
Hongcheng Gao, Hailong Qu, Jingyi Tang +18
Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world. However, existing benchmarks predomin…
REAL: A Reasoning-Enhanced Graph Framework for Long-Term Memory Management of LLMs
Keer Lu, Liwei Chen, Guoqing Jiang +3
Large Language Models (LLMs) are increasingly expected to interact with users over long time horizons. However, due to their finite context window, LLMs cannot retain all past inte…
MathMixup: Boosting LLM Mathematical Reasoning with Difficulty-Controllable Data Synthesis and Curriculum Learning
Xuchen Li, Jing Chen, Xuzhao Li +4
In mathematical reasoning tasks, the advancement of Large Language Models (LLMs) relies heavily on high-quality training data with clearly defined and well-graded difficulty levels…
SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning
Xuchen Li, Ruitao Wu, Xuanbo Liu +17
Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcr…
Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
Xuchen Li, Xuzhao Li, Jiahui Gao +3
Vision-Language Models (VLMs) excel at many multimodal tasks, yet they frequently struggle with tasks requiring precise understanding and handling of fine-grained visual elements.…
VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains
Xuzhao Li, Xuchen Li, Shiyu Hu +2
Large language models (LLMs) increasingly rely on reinforcement learning (RL) to enhance their reasoning capabilities through feedback. A critical challenge is verifying the consis…