1 paper
Sewade Ogun, Vincent Colotte, Emmanuel Vincent
Training of multi-speaker text-to-speech (TTS) systems relies on curated datasets based on high-quality recordings or audiobooks. Such datasets often lack speaker diversity and are…