3 papers
cs.SD2025
Discrete Audio Tokens: More Than a Survey!
Pooneh Mousavi, Gallil Maimon, Adel Moumen +18
Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and infere…
cs.CL2025
LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors
Rao Ma, Tongzhou Chen, Kartik Audhkhasi +1
Recently, large-scale pre-trained speech encoders and Large Language Models (LLMs) have been released, which show state-of-the-art performance on a range of spoken language process…
cs.CL2025
Schema Augmentation for Zero-Shot Domain Adaptation in Dialogue State Tracking
Christopher Richardson, Roshan Sharma, Neeraj Gaur +3
Zero-shot domain adaptation for dialogue state tracking (DST) remains a challenging problem in task-oriented dialogue (TOD) systems, where models must generalize to target domains…