109 citations · 131 across the 3 of their papers we have counts for
11 papers
MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation
Mingzhen Sun, Weining Wang, Yanyuan Qiao +5
Sounding Video Generation (SVG) is an audio-video joint generation task challenged by high-dimensional signal spaces, distinct data formats, and different patterns of content infor…
Calibration & Reconstruction: Deep Integrated Language for Referring Image Segmentation
Yichen Yan, Xingjian He, Sihan Chen +1
Referring image segmentation aims to segment an object referred to by natural language expression from an image. The primary challenge lies in the efficient propagation of fine-gra…
AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving Perception
Dingkang Yang, Shuai Huang, Zhi Xu +12
Driver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the…
MMNet: Multi-Mask Network for Referring Image Segmentation
Yichen Yan, Xingjian He, Wenxuan Wan +1
Referring image segmentation aims to segment an object referred to by natural language expression from an image. However, this task is challenging due to the distinct data properti…
Sounding Video Generator: A Unified Framework for Text-guided Sounding Video Generation
Jiawei Liu, Weining Wang, Sihan Chen +2
As a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames,…
MOSO: Decomposing MOtion, Scene and Object for Video Prediction
Mingzhen Sun, Weining Wang, Xinxin Zhu +1
Motion, scene and object are three primary visual components of a video. In particular, objects represent the foreground, scenes represent the background, and motion traces their d…