12 papers
Test-Time Compute Scaling for ASR with Depth-Conditioned Looped Transformers
Yacouba Kaloga, Shashi Kumar, Shakeel A. Sheikh +3
End-to-end ASR systems typically use fixed-depth acoustic encoders at inference, making it difficult to trade additional test-time computation for improved recognition without trai…
Geometric Latent Reasoning Induces Shorter Generations in LLMs
Shashi Kumar, Yacouba Kaloga, Petr Motlicek +2
Large language models solve complex problems by generating lengthy chains of explicit reasoning tokens. While effective, this makes reasoning expensive, length-sensitive, and const…
Data Augmentation for Pathological Speech Enhancement
Mingchi Hou, Enno Hermann, Ina Kodrasi
The performance of state-of-the-art speech enhancement (SE) models considerably degrades for pathological speech due to atypical acoustic characteristics and limited data availabil…
CLAP-Based Automatic Word Naming Recognition in Post-Stroke Aphasia
Yacouba Kaloga, Marina Laganaro, Ina Kodrasi
Conventional automatic word-naming recognition systems struggle to recognize words from post-stroke patients with aphasia because of disfluencies and mispronunciations, limiting re…
A Differentiable Alignment Framework for Sequence-to-Sequence Modeling via Optimal Transport
Yacouba Kaloga, Shashi Kumar, Petr Motlicek +1
Accurate sequence-to-sequence (seq2seq) alignment is critical for applications like medical speech analysis and language learning tools relying on automatic speech recognition (ASR…
Latent Space Factorization in LoRA
Shashi Kumar, Yacouba Kaloga, John Mitros +2
Low-rank adaptation (LoRA) is a widely used method for parameter-efficient finetuning. However, existing LoRA variants lack mechanisms to explicitly disambiguate task-relevant info…