9 papers
A French OSCE Dialogue Dataset and Controllable Virtual Patient System for Clinical Training
Doria Bonzi, Tom Bourgeade, Fabrice Lefèvre +1
The clinical and communication skills of medical students are commonly assessed through Objective Structured Clinical Examinations (OSCEs), which consist of brief scenario-driven s…
Synthetic Audio Generation Framework for Air Traffic Control Speech Recognition
Raphaël Bagat, Zhe Zhang, Junichi Yamagishi +2
Automatic Speech Recognition (ASR) systems, despite achieving remarkable accuracy in general-purpose domains with native speech (L1), struggle in domains like Air Traffic Control (…
BaldWhisper: Faster Whisper with Head Shearing and Layer Merging
Yaya Sy, Christophe Cerisara, Irina Illina
Pruning large pre-trained transformers in a data-scarce scenario is challenging, as it often requires massive retraining data to recover performance. For instance, Distill-Whisper…
Cross-lingual Matryoshka Representation Learning across Speech and Text
Yaya Sy, Dioula Doucouré, Christophe Cerisara +1
Speakers of under-represented languages face both a language barrier, as most online knowledge is in a few dominant languages, and a modality barrier, since information is largely…
LLM-Based Data Generation and Clinical Skills Evaluation for Low-Resource French OSCEs
Tian Huang, Tom Bourgeade, Irina Illina
Objective Structured Clinical Examinations (OSCEs) are the standard method for assessing medical students' clinical and communication skills through structured patient interviews.…
BEST-RQ-Based Self-Supervised Learning for Whisper Domain Adaptation
Raphaël Bagat, Irina Illina, Emmanuel Vincent
Automatic Speech Recognition (ASR) systems, despite large multilingual training, struggle in low-resource scenarios where labeled data is scarce. We propose BEARD (BEST-RQ Encoder…