3 papers
eess.AS2024
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
Emiru Tsunoo, Yuki Saito, Wataru Nakata +1
Real-time speech enhancement (SE) is essential to online speech communication. Causal SE models use only the previous context while predicting future information, such as phoneme c…
cs.SD2024
The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
Kaito Baba, Wataru Nakata, Yuki Saito +1
We present our system (denoted as T05) for the VoiceMOS Challenge (VMC) 2024. Our system was designed for the VMC 2024 Track 1, which focused on the accurate prediction of naturaln…
cs.SD2024
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
Kazuki Yamauchi, Yuki Saito, Hiroshi Saruwatari
We explore cross-dialect text-to-speech (CD-TTS), a task to synthesize learned speakers' voices in non-native dialects, especially in pitch-accent languages. CD-TTS is important fo…