16 papers
When Synthetic Speech Is All You Have: Better Call GRPO
Shashi Kumar, Yanis Labrak, Hasindri Watawana +5
LLM-based ASR adapted to regulated domains such as banking is bottlenecked by privacy: real speech is costly and legally constrained to collect, making synthetic text-to-speech (TT…
How to Leverage Synthetic Speech for LLM-Based ASR Systems?
Yanis Labrak, Dairazalia Sanchez-Cortes, Sergio Burdisso +9
In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic speech from modern text-to-speech (TTS) is a…
Evaluation of Automatic Speech Recognition Using Generative Large Language Models
Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil +6
Automatic Speech Recognition (ASR) is traditionally evaluated using Word Error Rate (WER), a metric that is insensitive to meaning. Embedding-based semantic metrics are better corr…
Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR
Thibault Bañeras-Roux, Sergio Burdisso, Esaú Villatoro-Tello +9
Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR architectures connect a speech encode…
Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization
Yanis Labrak, David Grünert, Séverin Baroudi +11
Long-context audio reasoning is underserved in both training data and evaluation. Existing benchmarks target short-context tasks, and the open-ended generation tasks most relevant…
Distilling Conversations: Abstract Compression of Conversational Audio Context for LLM-based ASR
Shashi Kumar, Esaú Villatoro-Tello, Sergio Burdisso +7
Standard LLM-based speech recognition systems typically process utterances in isolation, limiting their ability to leverage conversational context. In this work, we study whether m…