activity
20242026
most citedWeakly Supervised Detection and Temporal Localization of Whale Calls in Long-Duration Bioacoustic Data

1 citations · 1 across the 7 of their papers we have counts for

collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2026

Exploring Efficient Waveform Diffusion Models for Foley Sound Generation

Runwu Shi, Chang Li, Jiahui Li +7

Recent advances in diffusion models have enabled high-fidelity Foley sound generation directly in the waveform space. Existing waveform diffusion models primarily rely on time-doma…

eess.AS2026

Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS

Runwu Shi, Yujin Wang, Hongjin Song +1

Classifier-free guidance (CFG) is widely used in flow-matching-based zero-shot text-to-speech (TTS), where generation is typically controlled by two conditions: the target text and…

eess.AS2025

Unsupervised Single-Channel Audio Separation with Diffusion Source Priors

Runwu Shi, Chang Li, Jiang Wang +5

Single-channel audio separation aims to separate individual sources from a single-channel mixture. Most existing methods rely on supervised learning with synthetically generated pa…

eess.AS2025

Unsupervised Single-Channel Speech Separation with Diffusion under Speaker-Embedding Guidance

Runwu Shi, Kai Li, Chang Li +5

Speech separation is a fundamental task in audio processing, typically addressed with fully supervised systems trained on paired mixtures. While effective, such systems typically r…

eess.AS2025

Single-Channel Target Speech Extraction Utilizing Distance and Room Clues

Runwu Shi, Zirui Lin, Benjamin Yen +3

This paper aims to achieve single-channel target speech extraction (TSE) in enclosures utilizing distance clues and room information. Recent works have verified the feasibility of…

eess.AS2024

Bird Vocalization Embedding Extraction Using Self-Supervised Disentangled Representation Learning

Runwu Shi, Katsutoshi Itoyama, Kazuhiro Nakadai

This paper addresses the extraction of the bird vocalization embedding from the whole song level using disentangled representation learning (DRL). Bird vocalization embeddings are…