1 citations · 1 across the 3 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge
Yao Tang, Li Dong, Yaru Hao +3
Large language models often solve complex reasoning tasks more effectively with Chain-of-Thought (CoT), but at the cost of long, low-bandwidth token sequences. Humans, by contrast,…
cs.CL2025
Geometric-Mean Policy Optimization
Yuzhong Zhao, Yue Liu, Junpeng Liu +9
Group Relative Policy Optimization (GRPO) has significantly enhanced the reasoning capability of large language models by optimizing the arithmetic mean of token-level rewards. Unf…
cs.CL2025
Data Selection via Optimal Control for Language Models
Yuxian Gu, Li Dong, Hongning Wang +4
This work investigates the selection of high-quality pre-training data from massive corpora to enhance LMs' capabilities for downstream usage. We formulate data selection as a gene…