2 papers
cs.LG2026
Test-Time Compute Scaling for ASR with Depth-Conditioned Looped Transformers
Yacouba Kaloga, Shashi Kumar, Shakeel A. Sheikh +3
End-to-end ASR systems typically use fixed-depth acoustic encoders at inference, making it difficult to trade additional test-time computation for improved recognition without trai…
cs.CL2026
Evaluation of Automatic Speech Recognition Using Generative Large Language Models
Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil +6
Automatic Speech Recognition (ASR) is traditionally evaluated using Word Error Rate (WER), a metric that is insensitive to meaning. Embedding-based semantic metrics are better corr…