3 papers
cs.CL2024
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
Nick Rossenbach, Ralf Schlüter, Sakriani Sakti
The rapid development of neural text-to-speech (TTS) systems enabled its usage in other areas of natural language processing such as automatic speech recognition (ASR) or spoken la…
cs.CL2023
On the Relevance of Phoneme Duration Variability of Synthesized Training Data for Automatic Speech Recognition
Nick Rossenbach, Benedikt Hilmes, Ralf Schlüter
Synthetic data generated by text-to-speech (TTS) systems can be used to improve automatic speech recognition (ASR) systems in low-resource or domain mismatch tasks. It has been sho…
cs.CL2023
Take the Hint: Improving Arabic Diacritization with Partially-Diacritized Text
Parnia Bahar, Mattia Di Gangi, Nick Rossenbach +1
Automatic Arabic diacritization is useful in many applications, ranging from reading support for language learners to accurate pronunciation predictor for downstream tasks like spe…