2 papers
eess.AS2026
Chunkwise Aligners for Streaming Speech Recognition
Wen Shen Teo, Takafumi Moriya, Masato Mimura
We propose the Chunkwise Aligner, a novel architecture for streaming automatic speech recognition (ASR). While the Transducer is the standard model for streaming ASR, its training…
cs.CL2025
Task Arithmetic for Language Expansion in Speech Translation
Yao-Fei Cheng, Hayato Futami, Yosuke Kashiwagi +4
Recent progress in large language models (LLMs) has gained interest in speech-text multimodal foundation models, achieving strong performance on instruction-tuned speech translatio…