12 citations · 20 across the 5 of their papers we have counts for
5 papers
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
Jinxing Zhou, Dan Guo, Yuxin Mao +3
Audio-Visual Video Parsing (AVVP) task aims to detect and temporally locate events within audio and visual modalities. Multiple events can overlap in the timeline, making identific…
TAVGBench: Benchmarking Text to Audible-Video Generation
Yuxin Mao, Xuyang Shen, Jing Zhang +5
The Text to Audible-Video Generation (TAVG) task involves generating videos with accompanying audio based on text descriptions. Achieving this requires skillful alignment of both a…
Fine-grained Audible Video Description
Xuyang Shen, Dong Li, Jinxing Zhou +9
We explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audibl…
Improving Audio-Visual Video Parsing with Pseudo Visual Labels
Jinxing Zhou, Dan Guo, Yiran Zhong +1
Audio-Visual Video Parsing is a task to predict the events that occur in video segments for each modality. It often performs in a weakly supervised manner, where only video event l…
Audio-Visual Segmentation with Semantics
Jinxing Zhou, Xuyang Shen, Jianyuan Wang +8
We propose a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame…