2 papers
cs.SD2026
AugCodec: A Low-Bitrate Disentangled Neural Speech Codec via Data Augmentation
Dongmei Wang, Xiaohang Sun, Yang Liu +10
We propose AugCodec, a low-bitrate disentangled neural speech codec that leverages data augmentation to decompose speech into three distinct components: semantic, speaker, and pros…
eess.AS2025
Beyond Speaker Identity: Text Guided Target Speech Extraction
Mingyue Huo, Abhinav Jain, Cong Phuoc Huynh +4
Target Speech Extraction (TSE) traditionally relies on explicit clues about the speaker's identity like enrollment audio, face images, or videos, which may not always be available.…