collaborators

7 papers

cs.SD2026

Text-only adaptation in LLM-based ASR through text denoising

Andrés Carofilis, Sergio Burdisso, Esaú Villatoro-Tello +8

Adapting large language model (LLM)-based automatic speech recognition (ASR) systems to new domains using text-only data is a significant yet underexplored challenge. Standard fine…

cs.CL2025

TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation

Shashi Kumar, Srikanth Madikeri, Esaú Villatoro-Tello +8

Token-based multitasking frameworks like TokenVerse require all training utterances to have labels for all tasks, hindering their ability to leverage partially annotated datasets a…

cs.SD2025

Unifying Streaming and Non-streaming Zipformer-based ASR

Bidisha Sharma, Karthik Pandia Durai, Shankar Venkatesan +4

There has been increasing interest in unifying streaming and non-streaming automatic speech recognition (ASR) models to reduce development, training, and deployment costs. We prese…

eess.AS2025

Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM

Jeena Prakash, Blessingh Kumar, Kadri Hacioglu +5

Automatic speech recognition (ASR) models rely on high-quality transcribed data for effective training. Generating pseudo-labels for large unlabeled audio datasets often relies on…

cs.CL2025

Better Semi-supervised Learning for Multi-domain ASR Through Incremental Retraining and Data Filtering

Andres Carofilis, Pradeep Rangappa, Srikanth Madikeri +10

Fine-tuning pretrained ASR models for specific domains is challenging when labeled data is scarce. But unlabeled audio and labeled data from related domains are often available. We…

cs.CL2025

Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering

Pradeep Rangappa, Andres Carofilis, Jeena Prakash +10

Fine-tuning pretrained ASR models for specific domains is challenging for small organizations with limited labeled data and computational resources. Here, we explore different data…