activity
20242026
collaborators

23 papers

cs.CL2026

Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER

Hritika Sharma, Thibault Bañeras-Roux, Alessandra Pinto +4

Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equally costly, regardless of whet…

cs.CL2026

Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition

Hasindri Watawana, Sergio Burdisso, Esaú Villatoro-Tello +4

SpeechLLMs have shown strong potential for emotion recognition, yet they read the predicted emotion off a generative decoder not suited for classification: it can emit labels outsi…

cs.CL2026

Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study

Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil +6

Automatic Speech Recognition (ASR) is typically evaluated using Word Error Rate (WER), which poorly reflects semantic similarity. While embedding-based metrics correlate better wit…

cs.CL2026

When Synthetic Speech Is All You Have: Better Call GRPO

Shashi Kumar, Yanis Labrak, Hasindri Watawana +5

LLM-based ASR adapted to regulated domains such as banking is bottlenecked by privacy: real speech is costly and legally constrained to collect, making synthetic text-to-speech (TT…

cs.CL2026

How to Leverage Synthetic Speech for LLM-Based ASR Systems?

Yanis Labrak, Dairazalia Sanchez-Cortes, Sergio Burdisso +9

In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic speech from modern text-to-speech (TTS) is a…

cs.LG2026

Test-Time Compute Scaling for ASR with Depth-Conditioned Looped Transformers

Yacouba Kaloga, Shashi Kumar, Shakeel A. Sheikh +3

End-to-end ASR systems typically use fixed-depth acoustic encoders at inference, making it difficult to trade additional test-time computation for improved recognition without trai…