5 papers
Adaptive Supervised Anchoring for On-Policy Self-Distillation
Meilin Yang, Zixuan Ding, Jianhao Nie +5
On-policy self-distillation (OPSD) adapts a language model by distilling guidance from a frozen teacher on trajectories sampled from the student. Its effectiveness, however, depend…
MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model
Shanglin Yuan, Weiheng Zhao, Xianda Guo +4
Vision-language-action (VLA) models increasingly condition robot policies on history, depth, or 4D features to resolve ambiguity in long-horizon manipulation. However, more spatiot…
MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
Shuai Zhang, Bao Tang, Siyuan Yu +7
Recently, video generation has witnessed rapid advancements, drawing increasing attention to image-to-video (I2V) synthesis on mobile devices. However, the substantial computationa…
Image-Free Timestep Distillation via Continuous-Time Consistency with Trajectory-Sampled Pairs
Bao Tang, Shuai Zhang, Yueting Zhu +5
Timestep distillation is an effective approach for improving the generation efficiency of diffusion models. The Consistency Model (CM), as a trajectory-based framework, demonstrate…
LENS: Learning to Segment Anything with Unified Reinforced Reasoning
Lianghui Zhu, Bin Ouyang, Yuxuan Zhang +8
Text-prompted image segmentation enables fine-grained visual understanding and is critical for applications such as human-computer interaction and robotics. However, existing super…