5 papers · 1 filter
Magnitude-Direction Decoupling for Fast Video Generation with Flow Matching Models
Haonan Xu, Feiyang Chen, Songkui Chen +5
Flow matching models for video generation achieve impressive performance but suffer from high computational overhead due to iterative denoising. In fact, the original model is not…
Cross-Resolution Distribution Matching for Diffusion Distillation
Feiyang Chen, Hongpeng Pan, Haonan Xu +3
Diffusion distillation is central to accelerating image and video generation, yet existing methods are fundamentally limited by the denoising process, where step reduction has larg…
The Solution for CVPR2024 Foundational Few-Shot Object Detection Challenge
Hongpeng Pan, Shifeng Yi, Shouwei Yang +4
This report introduces an enhanced method for the Foundational Few-Shot Object Detection (FSOD) task, leveraging the vision-language model (VLM) for object detection. However, on s…
Learning to Rebalance Multi-Modal Optimization by Adaptively Masking Subnetworks
Yang Yang, Hongpeng Pan, Qing-Yuan Jiang +2
Multi-modal learning aims to enhance performance by unifying models from various modalities but often faces the "modality imbalance" problem in real data, leading to a bias towards…
Solution for Point Tracking Task of ICCV 1st Perception Test Challenge 2023
Hongpeng Pan, Yang Yang, Zhongtian Fu +4
This report proposes an improved method for the Tracking Any Point (TAP) task, which tracks any physical surface through a video. Several existing approaches have explored the TAP…