9 papers
Slot Filling as a Reasoning Task for SpeechLLMs
Kadri Hacioglu, Manjunath K E, Andreas Stolcke
We propose integration of reasoning into speech large language models (speechLLMs) for the end-to-end slot-filling task. Inspired by the recent development of reasoning LLMs, we us…
SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling
Kadri Hacioglu, Manjunath K E, Andreas Stolcke
Slot filling is a crucial subtask in spoken language understanding (SLU), traditionally implemented as a cascade of speech recognition followed by one or more natural language unde…
TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
Shashi Kumar, Srikanth Madikeri, Esaú Villatoro-Tello +8
Token-based multitasking frameworks like TokenVerse require all training utterances to have labels for all tasks, hindering their ability to leverage partially annotated datasets a…
Unifying Streaming and Non-streaming Zipformer-based ASR
Bidisha Sharma, Karthik Pandia Durai, Shankar Venkatesan +4
There has been increasing interest in unifying streaming and non-streaming automatic speech recognition (ASR) models to reduce development, training, and deployment costs. We prese…
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
Jeena Prakash, Blessingh Kumar, Kadri Hacioglu +5
Automatic speech recognition (ASR) models rely on high-quality transcribed data for effective training. Generating pseudo-labels for large unlabeled audio datasets often relies on…
Better Semi-supervised Learning for Multi-domain ASR Through Incremental Retraining and Data Filtering
Andres Carofilis, Pradeep Rangappa, Srikanth Madikeri +10
Fine-tuning pretrained ASR models for specific domains is challenging when labeled data is scarce. But unlabeled audio and labeled data from related domains are often available. We…