25 citations · 27 across the 3 of their papers we have counts for
3 papers
eess.AS2024
NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
Zhikang Niu, Sanyuan Chen, Long Zhou +3
Built upon vector quantization (VQ), discrete audio codec models have achieved great success in audio compression and auto-regressive audio generation. However, existing models fac…
cs.CL2023★ 25 cited
Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
Ziqiang Zhang, Long Zhou, Chengyi Wang +10
We propose a cross-lingual neural codec language model, VALL-E X, for cross-lingual speech synthesis. Specifically, we extend VALL-E and train a multi-lingual conditional codec lan…
eess.AS2021★ 2 cited
Self-Supervised Learning for speech recognition with Intermediate layer supervision
Chengyi Wang, Yu Wu, Sanyuan Chen +4
Recently, pioneer work finds that speech pre-trained models can solve full-stack speech processing tasks, because the model utilizes bottom layers to learn speaker-related informat…