11 papers
On-Policy Self-Distillation without Any Supervision
Yijiang Li, Bingyang Wang, Yijun Liang +3
On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). However, existing methods still rely heavily on external super…
Visual Contrastive Self-Distillation
Yijun Liang, Yunjie Tian, Yijiang Li +4
On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teach…
History-Conditioned Spatio-Temporal Visual Token Pruning for Efficient Vision-Language Navigation
Qitong Wang, Yijun Liang, Ming Li +2
Vision-Language Navigation (VLN) enables robots to follow natural-language instructions in visually grounded environments, serving as a key capability for embodied robotic systems.…
Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation
Lichen Li, Hengguang Zhou, Yijun Liang +2
Reward hacking in code generation, where models exploit evaluation loopholes to obtain high reward without correctly solving the intended task, poses a critical challenge for Reinf…
LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment
Han Chen, Ming Li, Chenguang Wang +4
Existing work on LLM-based educational assessment has focused largely on item difficulty, but difficulty alone does not indicate whether an item meaningfully distinguishes higher-…
Self-Evolving Visual Questioner
Yijun Liang, Hengguang Zhou, Ming Li +3
Vision-language models (VLMs) are typically trained as passive answerers, while their ability to actively ask diverse, non-trivial, visual-centric and grounded questions remains un…