3 papers
eess.AS2026
Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
Robert Flynn, Anton Ragni
Automatic speech recognition (ASR) models are normally trained to operate over single utterances, with a short duration of less than 30 seconds. This choice has been made in part d…
eess.AS2024
Self-Train Before You Transcribe
Robert Flynn, Anton Ragni
When there is a mismatch between the training and test domains, current speech recognition systems show significant performance degradation. Self-training methods, such as noisy st…
cs.CL2024
How Much Context Does My Attention-Based ASR System Need?
Robert Flynn, Anton Ragni
For the task of speech recognition, the use of more than 30 seconds of acoustic context during training is uncommon and under-investigated in literature. In this work, we conduct a…