2 citations · 2 across the 4 of their papers we have counts for
4 papers
TAVGBench: Benchmarking Text to Audible-Video Generation
Yuxin Mao, Xuyang Shen, Jing Zhang +5
The Text to Audible-Video Generation (TAVG) task involves generating videos with accompanying audio based on text descriptions. Achieving this requires skillful alignment of both a…
Multimodal Variational Auto-encoder based Audio-Visual Segmentation
Yuxin Mao, Jing Zhang, Mochu Xiang +2
We propose an Explicit Conditional Multimodal Variational Auto-Encoder (ECMVAE) for audio-visual segmentation (AVS), aiming to segment sound sources in the video sequence. Existing…
Measuring and Modeling Uncertainty Degree for Monocular Depth Estimation
Mochu Xiang, Jing Zhang, Nick Barnes +1
Effectively measuring and modeling the reliability of a trained model is essential to the real-world deployment of monocular depth estimation (MDE) models. However, the intrinsic i…
The Second Monocular Depth Estimation Challenge
Jaime Spencer, C. Stella Qian, Michaela Trescakova +40
This paper discusses the results for the second edition of the Monocular Depth Estimation Challenge (MDEC). This edition was open to methods using any form of supervision, includin…