2 papers
cs.CL2026
Merging the Knowledge of LLMs for Automatic Speech Recognition
Hayato Futami, Tatsuya Kawahara
Automatic speech recognition (ASR) systems, trained on paired speech-text data, have been improved by leveraging language models (LMs) trained on text-only data. LM fusion methods…
cs.SD2023
Zero- and Few-shot Sound Event Localization and Detection
Kazuki Shimada, Kengo Uchida, Yuichiro Koyama +4
Sound event localization and detection (SELD) systems estimate direction-of-arrival (DOA) and temporal activation for sets of target classes. Neural network (NN)-based SELD systems…