most citedBuboGPT: Enabling Visual Grounding in Multi-Modal LLMs

17 citations · 29 across the 10 of their papers we have counts for

collaborators

10 papers

cs.CV20241 cited

Real-time One-Step Diffusion-based Expressive Portrait Videos Generation

Hanzhong Guo, Hongwei Yi, Daquan Zhou +3

Latent diffusion models have made great strides in generating expressive portrait videos with accurate lip-sync and natural motion from a single reference image and audio input. Ho…

cs.CV20244 cited

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Lin Xu, Yilin Zhao, Daquan Zhou +3

Vision-language pre-training has significantly elevated performance across a wide range of image-language applications. Yet, the pre-training process for video-related tasks demand…

cs.CV20243 cited

Magic-Me: Identity-Specific Video Customized Diffusion

Ze Ma, Daquan Zhou, Chun-Hsiao Yeh +6

Creating content with specified identities (ID) has attracted significant interest in the field of generative models. In the field of text-to-image generation (T2I), subject-driven…

cs.CL20242 cited

Chain of Thought Explanation for Dialogue State Tracking

Lin Xu, Ningxin Peng, Daquan Zhou +2

Dialogue state tracking (DST) aims to record user queries and goals during a conversational interaction achieved by maintaining a predefined set of slots and their corresponding va…

cs.CV2024

Sora Generates Videos with Stunning Geometrical Consistency

Xuanyi Li, Daquan Zhou, Chenxu Zhang +3

The recently developed Sora model [1] has exhibited remarkable capabilities in video generation, sparking intense discussions regarding its ability to simulate real-world phenomena…

cs.CV20242 cited

MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation

Weimin Wang, Jiawei Liu, Zhijie Lin +9

The growing demand for high-fidelity video generation from textual descriptions has catalyzed significant research in this field. In this work, we introduce MagicVideo-V2 that inte…