9 papers · 1 filter
Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model
Xinyin Ma, Julius Berner, Chao Liu +3
Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods remain constrained by a rigid inference paradigm. Bidirectional d…
Q-ARVD: Quantizing Autoregressive Video Diffusion Models
Siao Tang, Xinyin Ma, Gongfan Fang +2
Autoregressive video diffusion models (ARVDs) have emerged as a promising architecture for streaming video generation, paving the way for real-time interactive video generation and…
Rethinking Token Reduction for Large Vision-Language Models
Yi Wang, Haofei Zhang, Qihan Huang +7
Large Vision-Language Models (LVLMs) excel in visual understanding and reasoning, but the excessive visual tokens lead to high inference costs. Although recent token reduction meth…
In-Video Instructions: Visual Signals as Generative Control
Gongfan Fang, Xinyin Ma, Xinchao Wang
Large-scale video generative models have recently demonstrated strong visual capabilities, enabling the prediction of future frames that adhere to the logical and physical cues in…
Diversity-Guided MLP Reduction for Efficient Large Vision Transformers
Chengchao Shen, Hourun Zhu, Gongfan Fang +2
Transformer models achieve excellent scaling property, where the performance is improved with the increment of model capacity. However, large-scale model parameters lead to an unaf…
PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud Learning
Song Wang, Xiaolu Liu, Lingdong Kong +6
Self-supervised representation learning for point cloud has demonstrated effectiveness in improving pre-trained model performance across diverse tasks. However, as pre-trained mode…