1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.SD2025
Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations
Yichen Han, Xiaoyang Hao, Keming Chen +25
Text-to-speech (TTS) synthesis has seen renewed progress under the discrete modeling paradigm. Existing autoregressive approaches often rely on single-codebook representations, whi…
cs.CL2025
Baichuan-Omni-1.5 Technical Report
Yadong Li, Jun Liu, Tao Zhang +89
We introduce Baichuan-Omni-1.5, an omni-modal model that not only has omni-modal understanding capabilities but also provides end-to-end audio generation capabilities. To achieve f…
cs.SD2025★ 1 cited
Towards Lightweight and Stable Zero-shot TTS with Self-distilled Representation Disentanglement
Qianniu Chen, Xiaoyang Hao, Bowen Li +2
Zero-shot Text-To-Speech (TTS) synthesis shows great promise for personalized voice customization through voice cloning. However, current methods for achieving zero-shot TTS heavil…