4 papers · 1 filter
Position: Towards Responsible Evaluation for Text-to-Speech
Yifan Yang, Hui Wang, Bing Han +4
Recent advances in text-to-speech (TTS) technology have enabled systems to generate speech that is often indistinguishable from human speech, bringing benefits to accessibility, co…
Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
Yifan Yang, Bing Han, Hui Wang +8
Modeling fine-grained speaking styles remains challenging for language-speech representation pre-training, as existing speech-text models are typically trained with coarse captions…
Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration
Yifan Yang, Bing Han, Hui Wang +5
Prosody diversity is essential for achieving naturalness and expressiveness in zero-shot text-to-speech (TTS). However, frequently used acoustic metrics capture only partial views…
Representation-Regularized Convolutional Audio Transformer for Audio Understanding
Bing Han, Chushu Zhou, Yifan Yang +4
Bootstrap-based Self-Supervised Learning (SSL) has achieved remarkable progress in audio understanding. However, existing methods typically operate at a single level of granularity…