2 papers
cs.LG2026
GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech
Antonis Asonitis, Francesco Verdini, Aref Farhadipour +4
We present GRAFT, a per-word pronunciation conditioning mechanism for text-to-speech neural codec language modeling. Existing systems reach high intelligibility and naturalness but…
eess.AS2024
An investigation of modularity for noise robustness in conformer-based ASR
Louise Coppieters de Gibson, Philip N. Garner, Pierre-Edouard Honnet
Whilst state of the art automatic speech recognition (ASR) can perform well, it still degrades when exposed to acoustic environments that differ from those used when training the m…