4 papers
Exploring Efficient Waveform Diffusion Models for Foley Sound Generation
Runwu Shi, Chang Li, Jiahui Li +7
Recent advances in diffusion models have enabled high-fidelity Foley sound generation directly in the waveform space. Existing waveform diffusion models primarily rely on time-doma…
PACE: Pretrained Audio Continual Learning
Chang Li, Kanglei Zhou, Liyuan Wang
Audio is a fundamental modality for analyzing speech, music, and environmental sounds. Although pretrained audio models have significantly advanced audio understanding, they remain…
Unsupervised Single-Channel Audio Separation with Diffusion Source Priors
Runwu Shi, Chang Li, Jiang Wang +5
Single-channel audio separation aims to separate individual sources from a single-channel mixture. Most existing methods rely on supervised learning with synthetically generated pa…
Unsupervised Single-Channel Speech Separation with Diffusion under Speaker-Embedding Guidance
Runwu Shi, Kai Li, Chang Li +5
Speech separation is a fundamental task in audio processing, typically addressed with fully supervised systems trained on paired mixtures. While effective, such systems typically r…