8 papers
Benchmarking Large Language Models for Grapheme-to-Phoneme Conversion: A Japanese Case Study
Tomoki Koriyama
Grapheme-to-phoneme (G2P) conversion is essential for controllable and robust text-to-speech, and large language models (LLMs), with broad linguistic knowledge, offer a promising a…
Exploring Pre-training Benefits on Phoneme Addition through Fine-tuning in Speech Synthesis
Masato Murata, Koichi Miyazaki, Tomoki Koriyama +1
Transfer learning is widely used for low-resource text-to-speech. When the target corpus contains phonemes unseen in pre-training, the model must expand its phoneme inventory durin…
Instantaneous Pitch Estimation via Wave-U-Net-Based Fundamental Waveform Enhancement
Junya Koguchi, Tomoki Koriyama
Instantaneous pitch estimation plays an important role in analyzing steep pitch variations such as speech prosody and singing techniques. Conventional approaches estimate instantan…
Voting-based Pitch Estimation with Temporal and Frequential Alignment and Correlation Aware Selection
Junya Koguchi, Tomoki Koriyama
The voting method, an ensemble approach for fundamental frequency estimation, is empirically known for its robustness but lacks thorough investigation. This paper provides a princi…
Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
Dong Yang, Yuki Saito, Takaaki Saeki +4
This paper advances phrase break prediction (also known as phrasing) in multi-speaker text-to-speech (TTS) systems. We integrate speaker-specific features by leveraging speaker emb…
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
Tomoki Koriyama
This paper proposes a model for automatic prosodic label annotation, where the predicted labels can be used for training a prosody-controllable text-to-speech model. The proposed m…