2 papers
cs.LG2026
VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-Distillation
Zewen Ding, Zezhong Wu, Zhou Tao +5
On-policy self-distillation (OPSD) improves reasoning by training a problem-only student on its own rollouts using dense token-level supervision from a privileged teacher that also…
cs.CV2026
MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction
Xiaokun Sun, Zezhong Wu, Zewen Ding +1
Reinforcement learning based post-training paradigms for Video Large Language Models (VideoLLMs) have achieved significant success by optimizing for visual-semantic tasks such as c…