4 papers
First Try Matters: Revisiting the Role of Reflection in Reasoning Models
Liwei Kang, Yue Deng, Yao Xiao +3
Large language models have recently demonstrated significant gains in reasoning ability, often attributed to their capacity to generate longer chains of thought and engage in refle…
Multi-Agent Tool-Integrated Policy Optimization
Zhanfeng Mo, Xingxuan Li, Yuntao Chen +1
Large language models (LLMs) increasingly rely on multi-turn tool-integrated planning for knowledge-intensive and complex reasoning tasks. Existing implementations typically rely o…
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization
Xingxuan Li, Yao Xiao, Dianwen Ng +15
Large language models have recently evolved from fluent text generation to advanced reasoning across diverse domains, giving rise to reasoning language models. Among these domains,…
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
Chong Zhang, Yue Deng, Xiang Lin +8
The recent development of reasoning language models (RLMs) represents a novel evolution in large language models. In particular, the recent release of DeepSeek-R1 has generated wid…