4 papers · 1 filter
LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors
Rao Ma, Tongzhou Chen, Kartik Audhkhasi +1
Recently, large-scale pre-trained speech encoders and Large Language Models (LLMs) have been released, which show state-of-the-art performance on a range of spoken language process…
Schema Augmentation for Zero-Shot Domain Adaptation in Dialogue State Tracking
Christopher Richardson, Roshan Sharma, Neeraj Gaur +3
Zero-shot domain adaptation for dialogue state tracking (DST) remains a challenging problem in task-oriented dialogue (TOD) systems, where models must generalize to target domains…
STAB: Speech Tokenizer Assessment Benchmark
Shikhar Vashishth, Harman Singh, Shikhar Bharadwaj +6
Representing speech as discrete tokens provides a framework for transforming speech into a format that closely resembles text, thus enabling the use of speech as an input to the wi…
Text Injection for Neural Contextual Biasing
Zhong Meng, Zelin Wu, Rohit Prabhavalkar +5
Neural contextual biasing effectively improves automatic speech recognition (ASR) for crucial phrases within a speaker's context, particularly those that are infrequent in the trai…