11 papers
MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models
Martino M. L. Pulici, Cuong Xuan Chu, Evgeny Kharlamov +3
Large language models achieve strong reasoning performance, but often at prohibitive training cost - a challenge that is especially acute for compact models (…
WorldReasoner: Evaluating Whether Language Model Agents Forecast Events with Valid Reasoning
Yizhou Chi, Eric Chamoun, Zifeng Ding +1
Forecasting real-world events requires language-model agents to reason under uncertainty from incomplete, time-bounded information. Yet evaluating whether agents genuinely forecast…
SciPaths: Forecasting Pathways to Scientific Discovery
Eric Chamoun, Yizhou Chi, Yulong Chen +4
Scientific progress depends on sequences of enabling contributions, yet existing AI4Science benchmarks largely focus on citation prediction, literature retrieval, or idea generatio…
Reasoning Compression with Mixed-Policy Distillation
Han Yang, Mingyan Wu, Bailan He +4
Reasoning-centric large language models (LLMs) achieve strong performance by generating intermediate reasoning trajectories, but often incur excessive token usage and high inferenc…
EigentSearch-Q+: Enhancing Deep Research Agents with Structured Reasoning Tools
Boer Zhang, Mingyan Wu, Dongzhuoran Zhou +6
Deep research requires reasoning over web evidence to answer open-ended questions, and it is a core capability for AI agents. Yet many deep research agents still rely on implicit,…
ProcessThinker: Enhancing Multi-modal Large Language Models Reasoning via Rollout-based Process Reward
Jingpei Wu, Xiao Han, Weixiang Shen +3
Visual question answering increasingly requires multi-step reasoning. Recent post-training with reinforcement learning under verifiable rewards (RLVR) and Group Relative Policy Opt…