From the 2 of 10 linked papers with an AI index.
10 papers
Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens
Daigo Takizawa, Tomohiko Nakamura, Samuele Cornell +3
The paper investigates how the language used to train neural audio codecs and self‑supervised speech models affects performance, finding that codec training language has little imp…
An Empirical Recipe for Universal Phone Recognition
Shikhar Bharadwaj, Chin-Jou Li, Kwanghee Choi +4
The paper introduces PhoneticXEUS, a multilingual phone recognition system trained on large-scale data that achieves state-of-the-art error rates on both many languages and accente…
ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era
Masao Someki, Alexander Polok, Carlos Carvalho +14
Recent speech research involves increasingly large datasets, complex models, and diverse experimental workflows. However, existing frameworks require substantial engineering effort…
Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption
Xun Gong, Jinchuan Tian, Haoran Wang +3
Current text-guided audio editing methods rely on paired training data, predefined operation templates, and separate processing pipelines across speech, music, and sound. We presen…
Evaluating Large Language Models Abilities for Addressee, Turn-change, and Next Speaker Prediction in Meetings
Ryo Fukuda, Takatomo Kano, Siddhant Arora +7
We investigate turn-taking in multimodal multi-party conversations using large language models (LLMs). We construct an evaluation framework for three tasks: addressee detection, tu…
Robust Training of Singing Voice Synthesis Using Prior and Posterior Uncertainty
Yiwen Zhao, Jiatong Shi, Yuxun Tang +2
Singing voice synthesis (SVS) has seen remarkable advancements in recent years. However, compared to speech and general audio data, publicly available singing datasets remain limit…