13 papers
The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding
Jiameng Zhang, Srikanth Madikeri
Large-scale video platforms process millions of uploads hourly, requiring moderation systems that can localize when and where policy violations occur within each video. Processing…
How to Leverage Synthetic Speech for LLM-Based ASR Systems?
Yanis Labrak, Dairazalia Sanchez-Cortes, Sergio Burdisso +9
In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic speech from modern text-to-speech (TTS) is a…
Distilling Conversations: Abstract Compression of Conversational Audio Context for LLM-based ASR
Shashi Kumar, Esaú Villatoro-Tello, Sergio Burdisso +7
Standard LLM-based speech recognition systems typically process utterances in isolation, limiting their ability to leverage conversational context. In this work, we study whether m…
Text-only adaptation in LLM-based ASR through text denoising
Andrés Carofilis, Sergio Burdisso, Esaú Villatoro-Tello +8
Adapting large language model (LLM)-based automatic speech recognition (ASR) systems to new domains using text-only data is a significant yet underexplored challenge. Standard fine…
TidyVoice 2026 Challenge Evaluation Plan
Aref Farhadipour, Jan Marquenie, Srikanth Madikeri +6
The performance of speaker verification systems degrades significantly under language mismatch, a critical challenge exacerbated by the field's reliance on English-centric data. To…
Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection
Sergio Burdisso, Esaú Villatoro-Tello, Shashi Kumar +7
LLM-based automatic speech recognition (ASR), a well-established approach, connects speech foundation models to large language models (LLMs) through a speech-to-LLM projector, yiel…