4 papers
High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
Chao Huang, Susan Liang, Yapeng Tian +2
We propose DAVIS, a Diffusion-based Audio-VIsual Separation framework that solves the audio-visual sound source separation task through generative learning. Existing methods typica…
FreSca: Scaling in Frequency Space Enhances Diffusion Models
Chao Huang, Susan Liang, Yunlong Tang +4
Latent diffusion models (LDMs) have achieved remarkable success in a variety of image tasks, yet achieving fine-grained, disentangled control over global structures versus fine det…
Language-Guided Joint Audio-Visual Editing via One-Shot Adaptation
Susan Liang, Chao Huang, Yapeng Tian +2
In this paper, we introduce a novel task called language-guided joint audio-visual editing. Given an audio and image pair of a sounding event, this task aims at generating new audi…
Scaling Concept With Text-Guided Diffusion Models
Chao Huang, Susan Liang, Yunlong Tang +3
Text-guided diffusion models have revolutionized generative tasks by producing high-fidelity content from text descriptions. They have also enabled an editing paradigm where concep…