10 papers · 1 filter
StreamTN: A Low-Latency Streaming Chinese Text Normalization Model for Streaming TTS in Dialogue Systems
Wenhao Li, Jinrui Liang, Haoyu Zhang +11
Text-to-Speech (TTS) is an essential module that provides spoken responses in a spoken dialogue system (SDS) centered on a large language model (LLM). To ensure accurate TTS synthe…
Rethinking Music Tokenization: A Semantic Codec toward High-Fidelity LLM Music Generation
Huakang Chen, Guobin Ma, Yuepeng Jiang +10
Discrete audio tokenization has become the critical interface between raw waveforms and autoregressive modeling in recent music generation. As a result, music tokenizers must simul…
SphereVAE: Hyperspherical Latent Autoencoders for Robust Autoregressive Speech Representation Modeling
Haoyu Zhang, Jingbin Hu, Hanke Xie +8
With the rapid development of speech generation technology, discrete codec representations have been widely used because they provide a stable prediction paradigm. In expressive sp…
SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation
Hanke Xie, Haopeng Lin, Jiale Qian +13
Continuous-latent autoregressive speech generation has emerged as a promising alternative to discrete-token modeling by avoiding quantization loss and preserving richer acoustic in…
MeanVC 2: Robust Low-Latency Streaming Zero-Shot Voice Conversion
Guobin Ma, Yuxuan Xia, Yuepeng Jiang +6
Streaming zero-shot voice conversion (VC) has become increasingly popular due to its potential for real-time applications. The recently proposed MeanVC achieves lightweight streami…
G-MaP-SE: Guided Speech Enhancement via GMM-Based Prior Matching
Yike Zhu, Ziqian Wang, Zikai Liu +5
Using speaker embeddings as conditioning can strengthen speech enhancement, but most methods either require clean enrollment audio or rely on embeddings extracted from noisy speech…