4 papers
GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech
Antonis Asonitis, Francesco Verdini, Aref Farhadipour +4
We present GRAFT, a per-word pronunciation conditioning mechanism for text-to-speech neural codec language modeling. Existing systems reach high intelligibility and naturalness but…
Post-Training Speech Enhancement Language Models with Perceptual Rewards
Frédéric Berdoz, Luca A. Lanzendörfer, Antonis Asonitis +1
Speech enhancement language models achieve strong results when trained on discrete audio tokens, but their optimization relies on token-level cross-entropy rather than the perceptu…
WorldSpeech: A Multilingual Speech Corpus from Around the World
Antonis Asonitis, Luca A. Lanzendörfer, Frédéric Berdoz +1
Automatic speech recognition (ASR) performs well for high-resource languages with abundant paired audio-transcript data, but its accuracy degrades sharply for most languages due to…
High-Fidelity Speech Enhancement via Discrete Audio Tokens
Luca A. Lanzendörfer, Frédéric Berdoz, Antonis Asonitis +1
Recent autoregressive transformer-based speech enhancement (SE) methods have shown promising results by leveraging advanced semantic understanding and contextual modeling of speech…