most citedA Comprehensive Analysis of Tokenization and Self-Supervised Learning in End-to-End Automatic Speech Recognition applied on French Language

1 citations · 1 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL20261 cited

A Comprehensive Analysis of Tokenization and Self-Supervised Learning in End-to-End Automatic Speech Recognition applied on French Language

Thibault Bañeras-Roux, Mickael Rouvier, Jane Wottawa +1

The performance of end-to-end automatic speech recognition (ASR) systems enables their increasing integration into numerous applications. While there are various benefits to such s…

cs.CL2026

A Paradigm for Interpreting Metrics and Identifying Critical Errors in Automatic Speech Recognition

Thibault Bañeras-Roux, Mickael Rouvier, Jane Wottawa +1

The most commonly used metrics for evaluating automatic speech transcriptions, namely Word Error Rate (WER) and Character Error Rate (CER), have been heavily criticized for their p…

cs.CL2026

MedInjection-FR: Exploring the Role of Native, Synthetic, and Translated Data in Biomedical Instruction Tuning

Ikram Belmadani, Oumaima El Khettari, Pacôme Constant dit Beaufils +2

Instruction tuning has become essential for adapting large language models (LLMs) to follow domain-specific prompts. Yet, in specialized fields such as medicine, the scarcity of hi…

cs.CL2025

An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training

Yanis Labrak, Richard Dufour, Mickaël Rouvier

This paper investigates discrete unit representations in Speech Language Models (SLMs), focusing on optimizing speech modeling during continual pre-training. In this paper, we syst…

cs.CL2025

Improving Social Determinants of Health Documentation in French EHRs Using Large Language Models

Adrien Bazoge, Pacôme Constant dit Beaufils, Mohammed Hmitouch +6

Social determinants of health (SDoH) significantly influence health outcomes, shaping disease progression, treatment adherence, and health disparities. However, their documentation…

cs.CL2025

A Benchmark of French ASR Systems Based on Error Severity

Antoine Tholly, Jane Wottawa, Mickael Rouvier +1

Automatic Speech Recognition (ASR) transcription errors are commonly assessed using metrics that compare them with a reference transcription, such as Word Error Rate (WER), which m…