5 papers
Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring
Hikaru Asano, Yotaro Kubo, So Kuroki
Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pro…
Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models
Yotaro Kubo, Qi Sun, Yujin Tang
This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic archi…
Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance
Manato Yaguchi, Yotaro Kubo, Hikaru Asano +1
Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying…
KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI
So Kuroki, Yotaro Kubo, Takuya Akiba +1
Real-time speech-to-speech (S2S) models excel at generating natural, low-latency conversational responses but often lack deep knowledge and semantic understanding. Conversely, casc…
Adaptive Dropout for Pruning Conformers
Yotaro Kubo, Xingyu Cai, Michiel Bacchiani
This paper proposes a method to effectively perform joint training-and-pruning based on adaptive dropout layers with unit-wise retention probabilities. The proposed method is based…