Showing eess.ASShow all
3 papers · 1 filter
eess.AS2025
Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
Dong Yang, Yuki Saito, Takaaki Saeki +4
This paper advances phrase break prediction (also known as phrasing) in multi-speaker text-to-speech (TTS) systems. We integrate speaker-specific features by leveraging speaker emb…
eess.AS2025
Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
Dong Yang, Yiyi Cai, Yuki Saito +2
We propose Shallow Flow Matching (SFM), a novel mechanism that enhances flow matching (FM)-based text-to-speech (TTS) models within a coarse-to-fine generation paradigm. Unlike con…
eess.AS2024
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
Emiru Tsunoo, Yuki Saito, Wataru Nakata +1
Real-time speech enhancement (SE) is essential to online speech communication. Causal SE models use only the previous context while predicting future information, such as phoneme c…