Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Streaming Translation and Transcription Through Speech-to-Text Causal Alignment
Roman Koshkin, Jeon Haesung, Lianbo Liu +4
Simultaneous machine translation (SiMT) has traditionally relied on offline machine translation models coupled with human-engineered heuristics or learned policies. We propose Hika…
cs.CL2026
DuplexCascade: Full-Duplex Speech-to-Speech Dialogue with VAD-Free Cascaded ASR-LLM-TTS Pipeline and Micro-Turn Optimization
Jianing Yang, Yusuke Fujita, Yui Sudo
Spoken dialog systems with cascaded ASR-LLM-TTS modules retain strong LLM intelligence, but VAD segmentation often forces half-duplex turns and brittle control. On the other hand,…
cs.CL2025
Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
Hao Shi, Yusuke Fujita, Tomoya Mizumoto +3
Prompts are crucial for task definition and for improving the performance of large language models (LLM)-based systems. However, existing LLM-based multi-talker (MT) automatic spee…