2 citations · 3 across the 11 of their papers we have counts for
8 papers · 1 filter
Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores
Jingjing Tang, Erica Cooper, Xin Wang +2
This paper presents an integrated system that transforms symbolic music scores into expressive piano performance audio. By combining a Transformer-based Expressive Performance Rend…
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
Wen-Chin Huang, Szu-Wei Fu, Erica Cooper +5
We present the third edition of the VoiceMOS Challenge, a scientific initiative designed to advance research into automatic prediction of human speech ratings. There were three tra…
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
Zhengyang Chen, Xuechen Liu, Erica Cooper +2
This paper proposes a speech synthesis system that allows users to specify and control the acoustic characteristics of a speaker by means of prompts describing the speaker's traits…
Uncertainty as a Predictor: Leveraging Self-Supervised Learning for Zero-Shot MOS Prediction
Aditya Ravuri, Erica Cooper, Junichi Yamagishi
Predicting audio quality in voice synthesis and conversion systems is a critical yet challenging task, especially when traditional methods like Mean Opinion Scores (MOS) are cumber…
Speaker-Text Retrieval via Contrastive Learning
Xuechen Liu, Xin Wang, Erica Cooper +2
In this study, we introduce a novel cross-modal retrieval task involving speaker descriptions and their corresponding audio samples. Utilizing pre-trained speaker and text encoders…
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
Cheng Gong, Xin Wang, Erica Cooper +5
Neural text-to-speech (TTS) has achieved human-like synthetic speech for single-speaker, single-language synthesis. Multilingual TTS systems are limited to resource-rich languages…