1 paper
Shengchuan Gao, Teng Hu, Bohao Feng +4
Recent audio-visual generation models can synthesize synchronized video and sound in a unified diffusion process, but their inference cost remains high because long video token seq…