most citedImitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

2 citations · 2 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CL2025

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR

Jia Deng, Jie Chen, Zhipeng Chen +7

Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs). Unlike traditiona…

cs.CL2025

Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning

Jia Deng, Jie Chen, Zhipeng Chen +2

Recently, reinforcement learning with verifiable rewards (RLVR) has been widely used for enhancing the reasoning abilities of large language models (LLMs). A core challenge in RLVR…

cs.CL2025

SimpleDeepSearcher: Deep Information Seeking via Web-Powered Reasoning Trajectory Synthesis

Shuang Sun, Huatong Song, Yuhao Wang +10

Retrieval-augmented generation (RAG) systems have advanced large language models (LLMs) in complex deep search scenarios requiring multi-step reasoning and iterative information re…

cs.CL2024

Enhancing LLM Reasoning with Reward-guided Tree Search

Jinhao Jiang, Zhipeng Chen, Yingqian Min +12

Recently, test-time scaling has garnered significant attention from the research community, largely due to the substantial advancements of the o1 model released by OpenAI. By alloc…

cs.CL2024

YuLan-Mini: An Open Data-efficient Language Model

Yiwen Hu, Huatong Song, Jia Deng +8

Effective pre-training of large language models (LLMs) has been challenging due to the immense resource demands and the complexity of the technical processes involved. This paper p…

cs.AI20242 cited

Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Yingqian Min, Zhipeng Chen, Jinhao Jiang +11

Recently, slow-thinking reasoning systems, such as o1, have demonstrated remarkable capabilities in solving complex reasoning tasks. These systems typically engage in an extended t…