1 paper · 1 filter
Kaichao Jiang, Changtao Miao, Baiqi Wu +9
Audio-video generation is rapidly moving from prompt-driven synthesis toward multimodal conditioning, where text, images, audio, and video can jointly shape the generated output. T…