activity
20242026
collaborators

5 papers

cs.LG2026

Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

Xinmu Ge, Zizhuo Zhang, Yu Huang +9

On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. It is commonly believed to enable the student model to distill knowledg…

cs.CL2026

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale

Ang Li, Ben Liu, Bin Han +215

Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong reasoning capabilities while remaining practical to train, serve,…

cs.LG2026

Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards

Yu Huang, Zihua Zhao, Zhaoxin Huan +9

The open-ended generation in LLMs usually requires multi-dimensional rubrics to adequately assess quality and guide the improvement of reinforcement learning. However, a critical d…

cs.LG2025

Bootstrapping your behavior: a new pretraining strategy for user behavior sequence data

Weichang Wu, Xiaolu Zhang, Jun Zhou +2

User Behavior Sequence (UBS) modeling is crucial in industrial applications. As data scale and task diversity grow, UBS pretraining methods have become increasingly pivotal. State-…

cs.LG2024

An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training

Youshao Xiao, Zhenglei Zhou, Fagui Mao +6

Recently, ChatGPT or InstructGPT like large language models (LLM) has made a significant impact in the AI world. Many works have attempted to reproduce the complex InstructGPT's tr…