5 papers
Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs?
Tomoya Mizumoto, Yusuke Fujita
Connecting a pre-trained speech encoder to a Large Language Model (LLM) is the standard architecture for building Speech LLMs. However, a structural misalignment exists between the…
Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language Models
Tomoya Mizumoto, Yusuke Fujita, Hao Shi +3
Dialogue systems based on large language models (LLMs) have advanced significantly in recent years. However, dialectal variation remains a major challenge, particularly for systems…
Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
Hao Shi, Yusuke Fujita, Tomoya Mizumoto +3
Prompts are crucial for task definition and for improving the performance of large language models (LLM)-based systems. However, existing LLM-based multi-talker (MT) automatic spee…
AC/DC: LLM-based Audio Comprehension via Dialogue Continuation
Yusuke Fujita, Tomoya Mizumoto, Atsushi Kojima +2
We propose an instruction-following audio comprehension model that leverages the dialogue continuation ability of large language models (LLMs). Instead of directly generating targe…
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
Yui Sudo, Yusuke Fujita, Atsushi Kojima +2
Speech foundation models (SFMs), such as Open Whisper-Style Speech Models (OWSM), are trained on massive datasets to achieve accurate automatic speech recognition. However, even SF…