most citedSimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

1 citations · 1 across the 1 of their papers we have counts for

collaborators

6 papers

cs.LG20251 cited

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Zhenghai Xue, Longtao Zheng, Qian Liu +4

Large Language Models (LLMs) can significantly improve their reasoning capabilities by interacting with external tools, a paradigm known as Tool-Integrated Reasoning (TIR). However…

cs.AI2025

First Return, Entropy-Eliciting Explore

Tianyu Zheng, Tianshun Xing, Qingshui Gu +10

Reinforcement Learning from Verifiable Rewards (RLVR) improves the reasoning abilities of Large Language Models (LLMs) but it struggles with unstable exploration. We propose FR3E (…

cs.LG2025

ZeCO: Zero Communication Overhead Sequence Parallelism for Linear Attention

Yuhong Chou, Zehao Liu, Ruijie Zhu +6

Linear attention mechanisms deliver significant advantages for Large Language Models (LLMs) by providing linear computational complexity, enabling efficient processing of ultra-lon…

cs.CL2025

General-Reasoner: Advancing LLM Reasoning Across All Domains

Xueguang Ma, Qian Liu, Dongfu Jiang +3

Reinforcement learning (RL) has recently demonstrated strong potential in enhancing the reasoning capabilities of large language models (LLMs). Particularly, the "Zero" reinforceme…

cs.SE2025

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

Mingzhe Du, Luu Anh Tuan, Yue Liu +6

Large Language Models (LLMs) generate functionally correct solutions but often fall short in code efficiency, a critical bottleneck for real-world deployment. In this paper, we int…

cs.LG2025

SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Weihao Zeng, Yuzhen Huang, Qian Liu +4

DeepSeek-R1 has shown that long chain-of-thought (CoT) reasoning can naturally emerge through a simple reinforcement learning (RL) framework with rule-based rewards, where the trai…