activity
20242026
collaborators

11 papers

cs.CL2026

When Synthetic Speech Is All You Have: Better Call GRPO

Shashi Kumar, Yanis Labrak, Hasindri Watawana +5

LLM-based ASR adapted to regulated domains such as banking is bottlenecked by privacy: real speech is costly and legally constrained to collect, making synthetic text-to-speech (TT…

cs.CL2026

How to Leverage Synthetic Speech for LLM-Based ASR Systems?

Yanis Labrak, Dairazalia Sanchez-Cortes, Sergio Burdisso +9

In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic speech from modern text-to-speech (TTS) is a…

cs.SD2026

Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization

Yanis Labrak, David Grünert, Séverin Baroudi +11

Long-context audio reasoning is underserved in both training data and evaluation. Existing benchmarks target short-context tasks, and the open-ended generation tasks most relevant…

eess.AS2026

Doctor or Patient? Synergizing Diarization and ASR for Code-Switched Hinglish Medical Conditions Extraction

Séverin Baroudi, Yanis Labrak, Shashi Kumar +7

Extracting patient medical conditions from code-switched clinical spoken dialogues is challenging due to rapid turn-taking and highly overlapped speech. We present a robust system…

cs.CL2025

LFAR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models

Santiago Cuervo, Adel Moumen, Yanis Labrak +5

Text-pretrained language models (LMs) encode rich world knowledge, but adapting them to process and generate perceptual modalities such as audio and images while effectively levera…

cs.CL2025

An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training

Yanis Labrak, Richard Dufour, Mickaël Rouvier

This paper investigates discrete unit representations in Speech Language Models (SLMs), focusing on optimizing speech modeling during continual pre-training. In this paper, we syst…