most citedCPTR: Full Transformer Network for Image Captioning

109 citations · 131 across the 3 of their papers we have counts for

collaborators

11 papers

cs.CV2024

MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation

Mingzhen Sun, Weining Wang, Yanyuan Qiao +5

Sounding Video Generation (SVG) is an audio-video joint generation task challenged by high-dimensional signal spaces, distinct data formats, and different patterns of content infor…

cs.CV20241 cited

Calibration & Reconstruction: Deep Integrated Language for Referring Image Segmentation

Yichen Yan, Xingjian He, Sihan Chen +1

Referring image segmentation aims to segment an object referred to by natural language expression from an image. The primary challenge lies in the efficient propagation of fine-gra…

cs.CV20231 cited

AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving Perception

Dingkang Yang, Shuai Huang, Zhi Xu +12

Driver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the…

cs.CV20235 cited

MMNet: Multi-Mask Network for Referring Image Segmentation

Yichen Yan, Xingjian He, Wenxuan Wan +1

Referring image segmentation aims to segment an object referred to by natural language expression from an image. However, this task is challenging due to the distinct data properti…

cs.CV202312 cited

Sounding Video Generator: A Unified Framework for Text-guided Sounding Video Generation

Jiawei Liu, Weining Wang, Sihan Chen +2

As a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames,…

cs.CV20231 cited

MOSO: Decomposing MOtion, Scene and Object for Video Prediction

Mingzhen Sun, Weining Wang, Xinxin Zhu +1

Motion, scene and object are three primary visual components of a video. In particular, objects represent the foreground, scenes represent the background, and motion traces their d…