1 citations · 1 across the 7 of their papers we have counts for
7 papers · 1 filter
Exploring Efficient Waveform Diffusion Models for Foley Sound Generation
Runwu Shi, Chang Li, Jiahui Li +7
Recent advances in diffusion models have enabled high-fidelity Foley sound generation directly in the waveform space. Existing waveform diffusion models primarily rely on time-doma…
Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS
Runwu Shi, Yujin Wang, Hongjin Song +1
Classifier-free guidance (CFG) is widely used in flow-matching-based zero-shot text-to-speech (TTS), where generation is typically controlled by two conditions: the target text and…
Unsupervised Single-Channel Audio Separation with Diffusion Source Priors
Runwu Shi, Chang Li, Jiang Wang +5
Single-channel audio separation aims to separate individual sources from a single-channel mixture. Most existing methods rely on supervised learning with synthetically generated pa…
Unsupervised Single-Channel Speech Separation with Diffusion under Speaker-Embedding Guidance
Runwu Shi, Kai Li, Chang Li +5
Speech separation is a fundamental task in audio processing, typically addressed with fully supervised systems trained on paired mixtures. While effective, such systems typically r…
Single-Channel Target Speech Extraction Utilizing Distance and Room Clues
Runwu Shi, Zirui Lin, Benjamin Yen +3
This paper aims to achieve single-channel target speech extraction (TSE) in enclosures utilizing distance clues and room information. Recent works have verified the feasibility of…
Bird Vocalization Embedding Extraction Using Self-Supervised Disentangled Representation Learning
Runwu Shi, Katsutoshi Itoyama, Kazuhiro Nakadai
This paper addresses the extraction of the bird vocalization embedding from the whole song level using disentangled representation learning (DRL). Bird vocalization embeddings are…