collaborators

6 papers

cs.SD2025

The AudioMOS Challenge 2025

Wen-Chin Huang, Hui Wang, Cheng Liu +6

This is the summary paper for the AudioMOS Challenge 2025, the very first challenge for automatic subjective quality prediction for synthetic audio. The challenge consists of three…

cs.SD2025

EchoVoices: Preserving Generational Voices and Memories for Seniors and Children

Haiying Xu, Haoze Liu, Mingshi Li +5

Recent breakthroughs in intelligent speech and digital human technologies have primarily targeted mainstream adult users, often overlooking the distinct vocal patterns and interact…

cs.SD2025

ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5

Jiaming Zhou, Shiyao Wang, Shiwan Zhao +10

Automatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However,…

eess.AS2024

Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge

Hongfei Xue, Rong Gong, Mingchen Shao +10

The StutteringSpeech Challenge focuses on advancing speech technologies for people who stutter, specifically targeting Stuttering Event Detection (SED) and Automatic Speech Recogni…

eess.AS2024

LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation

Wenhao Guan, Kaidi Wang, Wangjin Zhou +6

Recently, the application of diffusion models has facilitated the significant development of speech and audio generation. Nevertheless, the quality of samples generated by diffusio…

cs.SD2024

AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection

Rong Gong, Hongfei Xue, Lezhi Wang +11

The rapid advancements in speech technologies over the past two decades have led to human-level performance in tasks like automatic speech recognition (ASR) for fluent speech. Howe…