3 papers
cs.CL2025
SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling
Kadri Hacioglu, Manjunath K E, Andreas Stolcke
Slot filling is a crucial subtask in spoken language understanding (SLU), traditionally implemented as a cascade of speech recognition followed by one or more natural language unde…
cs.SD2025
Unifying Streaming and Non-streaming Zipformer-based ASR
Bidisha Sharma, Karthik Pandia Durai, Shankar Venkatesan +4
There has been increasing interest in unifying streaming and non-streaming automatic speech recognition (ASR) models to reduce development, training, and deployment costs. We prese…
cs.CL2024
Performance evaluation of SLAM-ASR: The Good, the Bad, the Ugly, and the Way Forward
Shashi Kumar, Iuliia Thorbecke, Sergio Burdisso +7
Recent research has demonstrated that training a linear connector between speech foundation encoders and large language models (LLMs) enables this architecture to achieve strong AS…