1 paper
Xin Yuan, Robin Feng, Mingming Ye
While deep learning-based text-to-speech (TTS) models such as VITS have shown excellent results, they typically require a sizable set of high-quality <text, audio> pairs to train,…