Publications (11)
Selective Cross-Task Distillation
Su Lu, Han-Jia Ye, De-Chuan Zhan
The outpouring of various pre-trained models empowers knowledge distillation by providing abundant teacher resources, but there lacks a developed mechanism to utilize these teacher…
Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training
Yuanyi Wang, Yifan Yang, Su Lu +9
Continual post-training aims to extend large language models (LLMs) with new knowledge, skills, and behaviors, yet it remains unclear when sequential updates enable capability tran…
InfiAlign: A Scalable and Sample-Efficient Framework for Aligning LLMs to Enhance Reasoning Capabilities
Shuo Cai, Su Lu, Qi Zhou +4
Large language models (LLMs) have exhibited impressive reasoning abilities on a wide range of complex tasks. However, enhancing these capabilities through post-training remains res…
Generalized Knowledge Distillation via Relationship Matching
Han-Jia Ye, Su Lu, De-Chuan Zhan
The knowledge of a well-trained deep neural network (a.k.a. the "teacher") is valuable for learning similar tasks. Knowledge distillation extracts knowledge from the teacher and in…
Industrial Dexterity Benchmark: A Hardware-Software Benchmarking Platform for Industrial Dexterous Manipulation
Honglu He, Jacob Laufer, Zhiwu Zheng +8
The paper introduces the Industrial Dexterity Benchmark (IDB) boards for evaluating industrial dexterous tasks, a scalable imitation‑learning framework (DAG‑ROS), and a multimodal…
Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation
Yuanyi Wang, Su Lu, Yanggan Gu +6
On-policy distillation (OPD) trains a student on its own rollouts with token-level teacher supervision. Recent selective OPD methods exploit the non-uniformity of OPD signals by pr…