most citedReward Reasoning Model

1 citations · 1 across the 3 of their papers we have counts for

collaborators

7 papers

cs.CL2026

Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge

Yao Tang, Li Dong, Yaru Hao +3

Large language models often solve complex reasoning tasks more effectively with Chain-of-Thought (CoT), but at the cost of long, low-bandwidth token sequences. Humans, by contrast,…

cs.AI2025

The Era of Agentic Organization: Learning to Organize with Language Models

Zewen Chi, Li Dong, Qingxiu Dong +4

We envision a new era of AI, termed agentic organization, where agents solve complex problems by working collaboratively and concurrently, enabling outcomes beyond individual intel…

cs.CL2025

Reinforcement Pre-Training

Qingxiu Dong, Li Dong, Yao Tang +4

In this work, we introduce Reinforcement Pre-Training (RPT) as a new scaling paradigm for large language models and reinforcement learning (RL). Specifically, we reframe next-token…

cs.CL2025

Think Only When You Need with Large Hybrid-Reasoning Models

Lingjie Jiang, Xun Wu, Shaohan Huang +7

Recent Large Reasoning Models (LRMs) have shown substantially improved reasoning capabilities over traditional Large Language Models (LLMs) by incorporating extended thinking proce…

cs.CL20251 cited

Reward Reasoning Model

Jiaxin Guo, Zewen Chi, Li Dong +4

Reward models play a critical role in guiding large language models toward outputs that align with human expectations. However, an open challenge remains in effectively utilizing t…

cs.CL2025

Scaling Laws of Synthetic Data for Language Models

Zeyu Qin, Qingxiu Dong, Xingxing Zhang +10

Large language models (LLMs) achieve strong performance across diverse tasks, largely driven by high-quality web data used in pre-training. However, recent studies indicate this da…