Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Video-MOPD: Multi-Teacher On-Policy Distillation for Video Understanding
Zhenxin Qin, Peng Shi, Cong Han +3
Video understanding demands a convergence of complementary capabilities across perception, temporal understanding, and complex reasoning, which are difficult to jointly optimize wi…
cs.CV2025
M4V: Multimodal Mamba for Efficient Text-to-Video Generation
Jiancheng Huang, Gengwei Zhang, Zequn Jie +5
Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators. However, modeling the vast spatiotemporal spa…
cs.CV2025
FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction
Siyu Jiao, Gengwei Zhang, Yinlong Qian +6
This work challenges the residual prediction paradigm in visual autoregressive modeling and presents FlexVAR, a new Flexible Visual AutoRegressive image generation paradigm. FlexVA…