2 papers
cs.CL2025
How I Built ASR for Endangered Languages with a Spoken Dictionary
Christopher Bartley, Anton Ragni
Nearly half of the world's languages are endangered. Speech technologies such as Automatic Speech Recognition (ASR) are central to revival efforts, yet most languages remain unsupp…
cs.CL2025
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video
Shumin Que, Anton Ragni
Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody given a single text input. While previous research has address…