2 papers
cs.SD2026
Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis
Lianbo Liu, Shiao Zhu, Kai Washizaki +10
While large language model (LLM)-based text-to-speech (TTS) systems have achieved high-quality speech synthesis, most existing systems focus on English and Chinese. Japanese, howev…
cs.CL2026
Streaming Translation and Transcription Through Speech-to-Text Causal Alignment
Roman Koshkin, Jeon Haesung, Lianbo Liu +4
Simultaneous machine translation (SiMT) has traditionally relied on offline machine translation models coupled with human-engineered heuristics or learned policies. We propose Hika…