collaborators

5 papers

cs.CL2026

How to Leverage Synthetic Speech for LLM-Based ASR Systems?

Yanis Labrak, Dairazalia Sanchez-Cortes, Sergio Burdisso +9

In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic speech from modern text-to-speech (TTS) is a…

cs.CL2026

Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR

Thibault Bañeras-Roux, Sergio Burdisso, Esaú Villatoro-Tello +9

Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR architectures connect a speech encode…

cs.SD2026

Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization

Yanis Labrak, David Grünert, Séverin Baroudi +11

Long-context audio reasoning is underserved in both training data and evaluation. Existing benchmarks target short-context tasks, and the open-ended generation tasks most relevant…

eess.AS2026

Doctor or Patient? Synergizing Diarization and ASR for Code-Switched Hinglish Medical Conditions Extraction

Séverin Baroudi, Yanis Labrak, Shashi Kumar +7

Extracting patient medical conditions from code-switched clinical spoken dialogues is challenging due to rapid turn-taking and highly overlapped speech. We present a robust system…

eess.AS2025

On the Use of Self-Supervised Representation Learning for Speaker Diarization and Separation

Séverin Baroudi, Hervé Bredin, Joseph Razik +1

Self-supervised speech models such as wav2vec2.0 and WavLM have been shown to significantly improve the performance of many downstream speech tasks, especially in low-resource sett…