6 papers
Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis
Lianbo Liu, Shiao Zhu, Kai Washizaki +10
While large language model (LLM)-based text-to-speech (TTS) systems have achieved high-quality speech synthesis, most existing systems focus on English and Chinese. Japanese, howev…
Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
Mengjie Zhao, Lianbo Liu, Yusuke Fujita +4
SpeechLLMs typically combine ASR-trained encoders with text-based LLM backbones, leading them to inherit written-style output patterns unsuitable for text-to-speech synthesis. This…
Streaming Translation and Transcription Through Speech-to-Text Causal Alignment
Roman Koshkin, Jeon Haesung, Lianbo Liu +4
Simultaneous machine translation (SiMT) has traditionally relied on offline machine translation models coupled with human-engineered heuristics or learned policies. We propose Hika…
Distilling LLM Semantic Priors into Encoder-Only Multi-Talker ASR with Talker-Count Routing
Hao Shi, Yusuke Fujita, Roman Koshkin +4
Large language models (LLMs) provide strong semantic priors that can improve multi-talker automatic speech recognition (MT-ASR), but using an LLM as an autoregressive decoder is co…
SASST: Leveraging Syntax-Aware Chunking and LLMs for Simultaneous Speech Translation
Zeyu Yang, Lai Wei, Roman Koshkin +2
This work proposes a grammar-based chunking strategy that segments input streams into semantically complete units by parsing dependency relations (e.g., noun phrase boundaries, ver…
MaRGen: Multi-Agent LLM Approach for Self-Directed Market Research and Analysis
Roman Koshkin, Pengyu Dai, Nozomi Fujikawa +2
We present an autonomous framework that leverages Large Language Models (LLMs) to automate end-to-end business analysis and market report generation. At its core, the system employ…