papers

Publications (11)

cs.LG2022

Selective Cross-Task Distillation

Su Lu, Han-Jia Ye, De-Chuan Zhan

The outpouring of various pre-trained models empowers knowledge distillation by providing abundant teacher resources, but there lacks a developed mechanism to utilize these teacher…

cs.LG2026

Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training

Yuanyi Wang, Yifan Yang, Su Lu +9

Continual post-training aims to extend large language models (LLMs) with new knowledge, skills, and behaviors, yet it remains unclear when sequential updates enable capability tran…

cs.AI2025

InfiAlign: A Scalable and Sample-Efficient Framework for Aligning LLMs to Enhance Reasoning Capabilities

Shuo Cai, Su Lu, Qi Zhou +4

Large language models (LLMs) have exhibited impressive reasoning abilities on a wide range of complex tasks. However, enhancing these capabilities through post-training remains res…

cs.CV2022

Generalized Knowledge Distillation via Relationship Matching

Han-Jia Ye, Su Lu, De-Chuan Zhan

The knowledge of a well-trained deep neural network (a.k.a. the "teacher") is valuable for learning similar tasks. Knowledge distillation extracts knowledge from the teacher and in…

cs.RO2026

Industrial Dexterity Benchmark: A Hardware-Software Benchmarking Platform for Industrial Dexterous Manipulation

Honglu He, Jacob Laufer, Zhiwu Zheng +8

The paper introduces the Industrial Dexterity Benchmark (IDB) boards for evaluating industrial dexterous tasks, a scalable imitation‑learning framework (DAG‑ROS), and a multimodal…

#industrial robotics#dexterous manipulation#imitation learning#multimodal perception
cs.LG2026

Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation

Yuanyi Wang, Su Lu, Yanggan Gu +6

On-policy distillation (OPD) trains a student on its own rollouts with token-level teacher supervision. Recent selective OPD methods exploit the non-uniformity of OPD signals by pr…