4 papers
Video-MOPD: Multi-Teacher On-Policy Distillation for Video Understanding
Zhenxin Qin, Peng Shi, Cong Han +3
Video understanding demands a convergence of complementary capabilities across perception, temporal understanding, and complex reasoning, which are difficult to jointly optimize wi…
On incremental and semi-global exponential stability of gradient flows satisfying generalized Łojasiewicz inequalities
Andreas Oliveira, Arthur C. B. de Oliveira, Mario Sznaier +1
The Łojasiewicz inequality characterizes objective-value convergence along gradient flows and, in special cases, yields exponential decay of the cost. However, such results do not…
M4V: Multimodal Mamba for Efficient Text-to-Video Generation
Jiancheng Huang, Gengwei Zhang, Zequn Jie +5
Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators. However, modeling the vast spatiotemporal spa…
FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction
Siyu Jiao, Gengwei Zhang, Yinlong Qian +6
This work challenges the residual prediction paradigm in visual autoregressive modeling and presents FlexVAR, a new Flexible Visual AutoRegressive image generation paradigm. FlexVA…