416 citations · 640 across the 20 of their papers we have counts for
7 papers · 1 filter
Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
Yifan Yang, Bing Han, Hui Wang +8
Modeling fine-grained speaking styles remains challenging for language-speech representation pre-training, as existing speech-text models are typically trained with coarse captions…
AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook
Yushen Chen, Kai Hu, Long Zhou +4
We propose AUV, a unified neural audio codec with a single codebook, which enables a favourable reconstruction of speech and further extends to general audio, including vocal, musi…
Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration
Yifan Yang, Bing Han, Hui Wang +5
Prosody diversity is essential for achieving naturalness and expressiveness in zero-shot text-to-speech (TTS). However, frequently used acoustic metrics capture only partial views…
WavChat: A Survey of Spoken Dialogue Models
Shengpeng Ji, Yifu Chen, Minghui Fang +16
Recent advancements in spoken dialogue models, exemplified by systems like GPT-4o, have captured significant attention in the speech domain. Compared to traditional three-tier casc…
NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
Zhikang Niu, Sanyuan Chen, Long Zhou +3
Built upon vector quantization (VQ), discrete audio codec models have achieved great success in audio compression and auto-regressive audio generation. However, existing models fac…
Robust Data2vec: Noise-robust Speech Representation Learning for ASR by Combining Regression and Improved Contrastive Learning
Qiu-Shi Zhu, Long Zhou, Jie Zhang +3
Self-supervised pre-training methods based on contrastive learning or regression tasks can utilize more unlabeled data to improve the performance of automatic speech recognition (A…