4 papers
TV-Regulated OPD: Direction Matters in On-Policy Distillation
Han Xiao, Yifan Niu, Dongyi Liu +2
On-Policy Distillation (OPD) facilitates the transfer of knowledge from domain expert to student in the post-training phase of Large Language Models (LLMs). However, the supervisio…
Breaking the Tokenizer Barrier: On-Policy Distillation across Model Families
Yifan Niu, Han Xiao, Dongyi Liu +4
On-Policy Distillation (OPD) has become a core technique in the post-training of Large Language Models (LLMs) for transferring knowledge from domain experts to student models. Howe…
PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning
Dongyi Liu, Yifan Niu, Qinwen Wang +2
Large Language Model (LLM)-based search agents trained with reinforcement learning (RL) have significantly improved the performance of knowledge-intensive tasks. However, existing…
Efficient Scaling of LLM Training with Flexible Context Parallelism
Yifan Niu, Han Xiao, Dongyi Liu +2
Scaling long-context capabilities is crucial for Large Language Models (LLMs). However, real-world data contain a large number of sequences with heterogeneous lengths. Existing tra…