most citedDual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source Localization

6 citations · 6 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV2025

TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation

Ziqian Wang, Yonghao He, Licheng Yang +6

Simulation provides a low-cost, scalable pathway to large-scale robotic manipulation data collection. However, existing 3D scene generation methods can rarely be applied directly t…

cs.CV2025

AudioStory: Generating Long-Form Narrative Audio with Large Language Models

Yuxin Guo, Teng Wang, Yuying Ge +4

Recent advances in text-to-audio (TTA) generation excel at synthesizing short audio clips but struggle with long-form narrative audio, which requires temporal coherence and composi…

cs.CV2025

Aligned Better, Listen Better for Audio-Visual Large Language Models

Yuxin Guo, Shuailei Ma, Shijie Ma +7

Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large la…

cs.CV20246 cited

Dual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source Localization

Yuxin Guo, Shijie Ma, Hu Su +5

Audio-Visual Source Localization (AVSL) aims to locate sounding objects within video frames given the paired audio clips. Existing methods predominantly rely on self-supervised con…

cs.CV2024

Cross Pseudo-Labeling for Semi-Supervised Audio-Visual Source Localization

Yuxin Guo, Shijie Ma, Yuhao Zhao +2

Audio-Visual Source Localization (AVSL) is the task of identifying specific sounding objects in the scene given audio cues. In our work, we focus on semi-supervised AVSL with pseud…