4 citations · 4 across the 9 of their papers we have counts for
1 paper · 1 filter
Le Xu, Chenxing Li, Yong Ren +5
Current vision-guided audio captioning systems frequently fail to address audiovisual misalignment in real-world scenarios, such as dubbed content or off-screen sounds. To bridge t…