1 paper · 1 filter
Pengjun Fang, Yingqing He, Yazhou Xing +3
Existing video-to-audio (V2A) generation methods predominantly rely on text prompts alongside visual information to synthesize audio. However, two critical bottlenecks persist: sem…