7 papers
Luna-TTS Family Technical Report
Feng Yin, Shuai Shi, Junjie Zheng +19
Modern text-to-speech (TTS) is dominated by autoregressive (AR) codec language models, whose left-to-right decoding brings latency that grows with utterance length, error accumulat…
Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions
Jinchuan Tian, Haoran Wang, Bo-Hao Su +14
Current audio foundation models typically rely on rigid, task-specific supervision (e.g., speech recognition), addressing isolated factors of audio rather than the whole. In contra…
Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption
Xun Gong, Jinchuan Tian, Haoran Wang +3
Current text-guided audio editing methods rely on paired training data, predefined operation templates, and separate processing pipelines across speech, music, and sound. We presen…
UrgentMOS: Unified Multi-Metric and Preference Learning for Robust Speech Quality Assessment
Wei Wang, Wangyou Zhang, Chenda Li +12
Automatic speech quality assessment has become increasingly important as modern speech generation systems continue to advance, while human listening tests remain costly, time-consu…
Text adaptation for speaker verification with speaker-text factorized embeddings
Yexin Yang, Shuai Wang, Xun Gong +2
Text mismatch between pre-collected data, either training data or enrollment data, and the actual test data can significantly hurt text-dependent speaker verification (SV) system p…
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
Xun Gong, Anqi Lv, Zhiming Wang +2
While speech large language models (SpeechLLMs) have advanced standard automatic speech recognition (ASR), contextual biasing for named entities and rare words remains challenging,…