most citedDuplexCascade: Full-Duplex Speech-to-Speech Dialogue with VAD-Free Cascaded ASR-LLM-TTS Pipeline and Micro-Turn Optimization

1 citations · 1 across the 6 of their papers we have counts for

collaborators

6 papers

cs.SD2026

Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization

Mengjie Zhao, Lianbo Liu, Yusuke Fujita +4

SpeechLLMs typically combine ASR-trained encoders with text-based LLM backbones, leading them to inherit written-style output patterns unsuitable for text-to-speech synthesis. This…

cs.SD2026

Distilling LLM Semantic Priors into Encoder-Only Multi-Talker ASR with Talker-Count Routing

Hao Shi, Yusuke Fujita, Roman Koshkin +4

Large language models (LLMs) provide strong semantic priors that can improve multi-talker automatic speech recognition (MT-ASR), but using an LLM as an autoregressive decoder is co…

cs.CL20261 cited

DuplexCascade: Full-Duplex Speech-to-Speech Dialogue with VAD-Free Cascaded ASR-LLM-TTS Pipeline and Micro-Turn Optimization

Jianing Yang, Yusuke Fujita, Yui Sudo

Spoken dialog systems with cascaded ASR-LLM-TTS modules retain strong LLM intelligence, but VAD segmentation often forces half-duplex turns and brittle control. On the other hand,…

cs.CL2025

Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition

Hao Shi, Yusuke Fujita, Tomoya Mizumoto +3

Prompts are crucial for task definition and for improving the performance of large language models (LLM)-based systems. However, existing LLM-based multi-talker (MT) automatic spee…

eess.AS2025

AC/DC: LLM-based Audio Comprehension via Dialogue Continuation

Yusuke Fujita, Tomoya Mizumoto, Atsushi Kojima +2

We propose an instruction-following audio comprehension model that leverages the dialogue continuation ability of large language models (LLMs). Instead of directly generating targe…

cs.SD2025

OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary

Yui Sudo, Yusuke Fujita, Atsushi Kojima +2

Speech foundation models (SFMs), such as Open Whisper-Style Speech Models (OWSM), are trained on massive datasets to achieve accurate automatic speech recognition. However, even SF…