1 paper
Gehui Chen, Guan'an Wang, Xiaowen Huang +1
Recent Video-to-Audio (V2A) generation relies on extracting semantic and temporal features from video to condition generative models. Training these models from scratch is resource…