3 papers
cs.SD2022
Improving Speech Prosody of Audiobook Text-to-Speech Synthesis with Acoustic and Textual Contexts
Detai Xin, Sharath Adavanne, Federico Ang +3
We present a multi-speaker Japanese audiobook text-to-speech (TTS) system that leverages multimodal context information of preceding acoustic context and bilateral textual context…
cs.SD2022
Mid-attribute speaker generation using optimal-transport-based interpolation of Gaussian mixture models
Aya Watanabe, Shinnosuke Takamichi, Yuki Saito +2
In this paper, we propose a method for intermediating multiple speakers' attributes and diversifying their voice characteristics in ``speaker generation,'' an emerging task that ai…
cs.SD2022
Speaking-Rate-Controllable HiFi-GAN Using Feature Interpolation
Detai Xin, Shinnosuke Takamichi, Takuma Okamoto +2
This paper presents a speaking-rate-controllable HiFi-GAN neural vocoder. Original HiFi-GAN is a high-fidelity, computationally efficient, and tiny-footprint neural vocoder. We att…