2 citations · 7 across the 6 of their papers we have counts for
6 papers
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
Jinxing Zhou, Dan Guo, Yuxin Mao +3
Audio-Visual Video Parsing (AVVP) task aims to detect and temporally locate events within audio and visual modalities. Multiple events can overlap in the timeline, making identific…
TAVGBench: Benchmarking Text to Audible-Video Generation
Yuxin Mao, Xuyang Shen, Jing Zhang +5
The Text to Audible-Video Generation (TAVG) task involves generating videos with accompanying audio based on text descriptions. Achieving this requires skillful alignment of both a…
Multimodal Variational Auto-encoder based Audio-Visual Segmentation
Yuxin Mao, Jing Zhang, Mochu Xiang +2
We propose an Explicit Conditional Multimodal Variational Auto-Encoder (ECMVAE) for audio-visual segmentation (AVS), aiming to segment sound sources in the video sequence. Existing…
RPEFlow: Multimodal Fusion of RGB-PointCloud-Event for Joint Optical Flow and Scene Flow Estimation
Zhexiong Wan, Yuxin Mao, Jing Zhang +1
Recently, the RGB images and point clouds fusion methods have been proposed to jointly estimate 2D optical flow and 3D scene flow. However, as both conventional RGB cameras and LiD…
Decomposed Guided Dynamic Filters for Efficient RGB-Guided Depth Completion
Yufei Wang, Yuxin Mao, Qi Liu +1
RGB-guided depth completion aims at predicting dense depth maps from sparse depth measurements and corresponding RGB images, where how to effectively and efficiently exploit the mu…
Mutual Information Regularization for Weakly-supervised RGB-D Salient Object Detection
Aixuan Li, Yuxin Mao, Jing Zhang +1
In this paper, we present a weakly-supervised RGB-D salient object detection model via scribble supervision. Specifically, as a multimodal learning task, we focus on effective mult…