1 paper
Xiulong Liu, Kun Su, Eli Shlizerman
The content of visual and audio scenes is multi-faceted such that a video can be paired with various audio and vice-versa. Thereby, in video-to-audio generation task, it is imperat…