4 papers · 1 filter
Magnitude-Direction Decoupling for Fast Video Generation with Flow Matching Models
Haonan Xu, Feiyang Chen, Songkui Chen +5
Flow matching models for video generation achieve impressive performance but suffer from high computational overhead due to iterative denoising. In fact, the original model is not…
Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini
Madhuri Shanbhogue, Zhe Li, Shanfeng Zhang +86
We introduce Gemini Embedding 2, a native multimodal embedding model that allows embedding video, audio, image, and text modalities in a unified representation space. We leverage t…
Cross-Resolution Distribution Matching for Diffusion Distillation
Feiyang Chen, Hongpeng Pan, Haonan Xu +3
Diffusion distillation is central to accelerating image and video generation, yet existing methods are fundamentally limited by the denoising process, where step reduction has larg…
Comprehensive Survey of Model Compression and Speed up for Vision Transformers
Feiyang Chen, Ziqian Luo, Lisang Zhou +2
Vision Transformers (ViT) have marked a paradigm shift in computer vision, outperforming state-of-the-art models across diverse tasks. However, their practical deployment is hamper…