1 paper · 1 filter
Junwon Lee, Juhan Nam, Jiyoung Lee
This work introduces a new task, text-conditioned selective video-to-audio (V2A) generation, which produces only the user-intended sound from a multi-object video. This capability…