2 citations · 4 across the 6 of their papers we have counts for
4 papers · 1 filter
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding
Yang Yang, Yunpeng Li, George Sung +4
Token-based language modeling is a prominent approach for speech generation, where tokens are obtained by quantizing features from self-supervised learning (SSL) models and extract…
Binaural Angular Separation Network
Yang Yang, George Sung, Shao-Fu Shih +3
We propose a neural network model that can separate target speech sources from interfering sources at different angular regions using two microphones. The model is trained with sim…
StreamVC: Real-Time Low-Latency Voice Conversion
Yang Yang, Yury Kartynnik, Yunpeng Li +4
We present StreamVC, a streaming voice conversion solution that preserves the content and prosody of any source speech while matching the voice timbre from any target speech. Unlik…
Guided Speech Enhancement Network
Yang Yang, Shao-Fu Shih, Hakan Erdogan +5
High quality speech capture has been widely studied for both voice communication and human computer interface reasons. To improve the capture performance, we can often find multi-m…