activity
20242026
most citedMultimodal Latent Language Modeling with Next-Token Diffusion

1 citations · 1 across the 6 of their papers we have counts for

collaborators

7 papers

cs.CL2026

Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge

Yao Tang, Li Dong, Yaru Hao +3

Large language models often solve complex reasoning tasks more effectively with Chain-of-Thought (CoT), but at the cost of long, low-bandwidth token sequences. Humans, by contrast,…

cs.LG2025

BitNet Distillation

Xun Wu, Shaohan Huang, Wenhui Wang +4

In this paper, we present BitNet Distillation (BitDistill), a lightweight pipeline that fine-tunes off-the-shelf full-precision LLMs (e.g., Qwen) into 1.58-bit precision (i.e., ter…

cs.CL2025

Information-Preserving Reformulation of Reasoning Traces for Antidistillation

Jiayu Ding, Lei Cui, Li Dong +2

Recent advances in Large Language Models (LLMs) show that extending the length of reasoning chains significantly improves performance on complex tasks. While revealing these reason…

cs.CL2025

Reinforcement Pre-Training

Qingxiu Dong, Li Dong, Yao Tang +4

In this work, we introduce Reinforcement Pre-Training (RPT) as a new scaling paradigm for large language models and reinforcement learning (RL). Specifically, we reframe next-token…

cs.CL2025

Rectified Sparse Attention

Yutao Sun, Tianzhu Ye, Li Dong +6

Efficient long-sequence generation is a critical challenge for Large Language Models. While recent sparse decoding methods improve efficiency, they suffer from KV cache misalignmen…

cs.CL2025

Scaling Laws of Synthetic Data for Language Models

Zeyu Qin, Qingxiu Dong, Xingxing Zhang +10

Large language models (LLMs) achieve strong performance across diverse tasks, largely driven by high-quality web data used in pre-training. However, recent studies indicate this da…