Showing eess.ASShow all
3 papers · 1 filter
eess.AS2025
Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model in End-to-End Speech Recognition
Yosuke Higuchi, Tetsuji Ogawa, Tetsunori Kobayashi
We propose to utilize an instruction-tuned large language model (LLM) for guiding the text generation process in automatic speech recognition (ASR). Modern large language models (L…
eess.AS2024
End-to-End Speech Recognition with Pre-trained Masked Language Model
Yosuke Higuchi, Tetsuji Ogawa, Tetsunori Kobayashi +1
We present a novel approach to end-to-end automatic speech recognition (ASR) that utilizes pre-trained masked language models (LMs) to facilitate the extraction of linguistic infor…
eess.AS2024
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
Oswald Zink, Yosuke Higuchi, Carlos Mullov +2
Effective spoken dialog systems should facilitate natural interactions with quick and rhythmic timing, mirroring human communication patterns. To reduce response times, previous ef…