17 citations · 29 across the 10 of their papers we have counts for
10 papers
Real-time One-Step Diffusion-based Expressive Portrait Videos Generation
Hanzhong Guo, Hongwei Yi, Daquan Zhou +3
Latent diffusion models have made great strides in generating expressive portrait videos with accurate lip-sync and natural motion from a single reference image and audio input. Ho…
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Lin Xu, Yilin Zhao, Daquan Zhou +3
Vision-language pre-training has significantly elevated performance across a wide range of image-language applications. Yet, the pre-training process for video-related tasks demand…
Magic-Me: Identity-Specific Video Customized Diffusion
Ze Ma, Daquan Zhou, Chun-Hsiao Yeh +6
Creating content with specified identities (ID) has attracted significant interest in the field of generative models. In the field of text-to-image generation (T2I), subject-driven…
Chain of Thought Explanation for Dialogue State Tracking
Lin Xu, Ningxin Peng, Daquan Zhou +2
Dialogue state tracking (DST) aims to record user queries and goals during a conversational interaction achieved by maintaining a predefined set of slots and their corresponding va…
Sora Generates Videos with Stunning Geometrical Consistency
Xuanyi Li, Daquan Zhou, Chenxu Zhang +3
The recently developed Sora model [1] has exhibited remarkable capabilities in video generation, sparking intense discussions regarding its ability to simulate real-world phenomena…
MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation
Weimin Wang, Jiawei Liu, Zhijie Lin +9
The growing demand for high-fidelity video generation from textual descriptions has catalyzed significant research in this field. In this work, we introduce MagicVideo-V2 that inte…