papers

Publications (50)

eess.AS2022

Generalization Ability of MOS Prediction Networks

Erica Cooper, Wen-Chin Huang, Tomoki Toda +1

Automatic methods to predict listener opinions of synthesized speech remain elusive since listeners, systems being evaluated, characteristics of the speech, and even the instructio…

eess.AS2024

Spoofing-Aware Speaker Verification Robust Against Domain and Channel Mismatches

Chang Zeng, Xiaoxiao Miao, Xin Wang +2

In real-world applications, it is challenging to build a speaker verification system that is simultaneously robust against common threats, including spoofing attacks, channel misma…

cs.SD2021

LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech

Wen-Chin Huang, Erica Cooper, Junichi Yamagishi +1

An effective approach to automatically predict the subjective rating for synthetic speech is to train on a listening test dataset with human-annotated scores. Although each speech…

eess.AS2023

Investigating Range-Equalizing Bias in Mean Opinion Score Ratings of Synthesized Speech

Erica Cooper, Junichi Yamagishi

Mean Opinion Score (MOS) is a popular measure for evaluating synthesized speech. However, the scores obtained in MOS tests are heavily dependent upon many contextual factors. One s…

eess.AS2020

How Similar or Different Is Rakugo Speech Synthesizer to Professional Performers?

Shuhei Kato, Yusuke Yasuda, Xin Wang +2

We have been working on speech synthesis for rakugo (a traditional Japanese form of verbal entertainment similar to one-person stand-up comedy) toward speech synthesis that authent…

eess.AS2021

Use of speaker recognition approaches for learning and evaluating embedding representations of musical instrument sounds

Xuan Shi, Erica Cooper, Junichi Yamagishi

Constructing an embedding space for musical instrument sounds that can meaningfully represent new and unseen instruments is important for downstream music generation tasks such as…

cs.SD2020

Pretraining Strategies, Waveform Model Choice, and Acoustic Configurations for Multi-Speaker End-to-End Speech Synthesis

Erica Cooper, Xin Wang, Yi Zhao +2

We explore pretraining strategies including choice of base corpus with the aim of choosing the best strategy for zero-shot multi-speaker end-to-end synthesis. We also examine choic…

eess.AS2020

Modeling of Rakugo Speech and Its Limitations: Toward Speech Synthesis That Entertains Audiences

Shuhei Kato, Yusuke Yasuda, Xin Wang +3

We have been investigating rakugo speech synthesis as a challenging example of speech synthesis that entertains audiences. Rakugo is a traditional Japanese form of verbal entertain…

cs.SD2023

Can Knowledge of End-to-End Text-to-Speech Models Improve Neural MIDI-to-Audio Synthesis Systems?

Xuan Shi, Erica Cooper, Xin Wang +2

With the similarity between music and speech synthesis from symbolic input and the rapid development of text-to-speech (TTS) techniques, it is worthwhile to explore ways to improve…

eess.AS2020

Zero-Shot Multi-Speaker Text-To-Speech with State-of-the-art Neural Speaker Embeddings

Erica Cooper, Cheng-I Lai, Yusuke Yasuda +4

While speaker adaptation for end-to-end speech synthesis using speaker embeddings can produce good speaker similarity for speakers seen during training, there remains a gap for zer…

eess.AS2026

MOS-Bias: From Hidden Gender Bias to Gender-Aware Speech Quality Assessment

Wenze Ren, Yi-Cheng Lin, Wen-Chin Huang +5

The Mean Opinion Score (MOS) serves as the standard metric for speech quality assessment, yet biases in human annotations remain underexplored. We conduct the first systematic anal…

cs.SD2024

Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems

Zhengyang Chen, Xuechen Liu, Erica Cooper +2

This paper proposes a speech synthesis system that allows users to specify and control the acoustic characteristics of a speaker by means of prompts describing the speaker's traits…

cs.SD2022

Analyzing Language-Independent Speaker Anonymization Framework under Unseen Conditions

Xiaoxiao Miao, Xin Wang, Erica Cooper +2

In our previous work, we proposed a language-independent speaker anonymization system based on self-supervised learning models. Although the system can anonymize speech data of any…

cs.SD2023

Speaker-Text Retrieval via Contrastive Learning

Xuechen Liu, Xin Wang, Erica Cooper +2

In this study, we introduce a novel cross-modal retrieval task involving speaker descriptions and their corresponding audio samples. Utilizing pre-trained speaker and text encoders…

cs.SD2026

MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models

Wen-Chin Huang, Erica Cooper, Tomoki Toda

In this paper, we study the task of subjective speech quality assessment (SSQA), which refers to predicting the perceptual quality of speech. Owing to the development of deep neura…

cs.SD2023

DDSP-based Neural Waveform Synthesis of Polyphonic Guitar Performance from String-wise MIDI Input

Nicolas Jonason, Xin Wang, Erica Cooper +3

We explore the use of neural synthesis for acoustic guitar from string-wise MIDI input. We propose four different systems and compare them with both objective metrics and subjectiv…

cs.SD2025

The AudioMOS Challenge 2025

Wen-Chin Huang, Hui Wang, Cheng Liu +6

This is the summary paper for the AudioMOS Challenge 2025, the very first challenge for automatic subjective quality prediction for synthetic audio. The challenge consists of three…

cs.SD2022

Text-to-Speech Synthesis Techniques for MIDI-to-Audio Synthesis

Erica Cooper, Xin Wang, Junichi Yamagishi

Speech synthesis and music audio generation from symbolic input differ in many aspects but share some similarities. In this study, we investigate how text-to-speech synthesis techn…

cs.SD2021

On the Interplay Between Sparsity, Naturalness, Intelligibility, and Prosody in Speech Synthesis

Cheng-I Jeff Lai, Erica Cooper, Yang Zhang +8

Are end-to-end text-to-speech (TTS) models over-parametrized? To what extent can these models be pruned, and what happens to their synthesis capabilities? This work serves as a sta…

eess.AS2025

HighRateMOS: Sampling-Rate Aware Modeling for Speech Quality Assessment

Wenze Ren, Yi-Cheng Lin, Wen-Chin Huang +9

Modern speech quality prediction models are trained on audio data resampled to a specific sampling rate. When faced with higher-rate audio at test time, these models can produce bi…

cs.SD2023

Uncertainty as a Predictor: Leveraging Self-Supervised Learning for Zero-Shot MOS Prediction

Aditya Ravuri, Erica Cooper, Junichi Yamagishi

Predicting audio quality in voice synthesis and conversion systems is a critical yet challenging task, especially when traditional methods like Mean Opinion Scores (MOS) are cumber…

cs.SD2023

Speaker anonymization using orthogonal Householder neural network

Xiaoxiao Miao, Xin Wang, Erica Cooper +2

Speaker anonymization aims to conceal a speaker's identity while preserving content information in speech. Current mainstream neural-network speaker anonymization systems disentang…

cs.SD2023

Range-Based Equal Error Rate for Spoof Localization

Lin Zhang, Xin Wang, Erica Cooper +2

Spoof localization, also called segment-level detection, is a crucial task that aims to locate spoofs in partially spoofed audio. The equal error rate (EER) is widely used to measu…

cs.SD2024

The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction

Wen-Chin Huang, Szu-Wei Fu, Erica Cooper +5

We present the third edition of the VoiceMOS Challenge, a scientific initiative designed to advance research into automatic prediction of human speech ratings. There were three tra…

cs.SD2024

ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations

Cheng Gong, Xin Wang, Erica Cooper +5

Neural text-to-speech (TTS) has achieved human-like synthetic speech for single-speaker, single-language synthesis. Multilingual TTS systems are limited to resource-rich languages…

eess.AS2021

Learning Disentangled Phone and Speaker Representations in a Semi-Supervised VQ-VAE Paradigm

Jennifer Williams, Yi Zhao, Erica Cooper +1

We present a new approach to disentangle speaker voice and phone content by introducing new components to the VQ-VAE architecture for speech synthesis. The original VQ-VAE does not…

cs.CL2024

An Initial Investigation of Language Adaptation for TTS Systems under Low-resource Scenarios

Cheng Gong, Erica Cooper, Xin Wang +9

Self-supervised learning (SSL) representations from massively multilingual models offer a promising solution for low-resource language speech tasks. Despite advancements, language…

cs.SD2021

Multi-Task Learning in Utterance-Level and Segmental-Level Spoof Detection

Lin Zhang, Xin Wang, Erica Cooper +1

In this paper, we provide a series of multi-tasking benchmarks for simultaneously detecting spoofing at the segmental and utterance levels in the PartialSpoof database. First, we p…

eess.AS2023

Improving Generalization Ability of Countermeasures for New Mismatch Scenario by Combining Multiple Advanced Regularization Terms

Chang Zeng, Xin Wang, Xiaoxiao Miao +2

The ability of countermeasure models to generalize from seen speech synthesis methods to unseen ones has been investigated in the ASVspoof challenge. However, a new mismatch scenar…

eess.AS2020

Can Speaker Augmentation Improve Multi-Speaker End-to-End TTS?

Erica Cooper, Cheng-I Lai, Yusuke Yasuda +1

Previous work on speaker adaptation for end-to-end speech synthesis still falls short in speaker similarity. We investigate an orthogonal approach to the current speaker adaptation…

cs.SD2022

The VoiceMOS Challenge 2022

Wen-Chin Huang, Erica Cooper, Yu Tsao +3

We present the first edition of the VoiceMOS Challenge, a scientific event that aims to promote the study of automatic prediction of the mean opinion score (MOS) of synthetic speec…

cs.SD2026

CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents

Wen-Chin Huang, Nicholas Sanders, Erica Cooper

We present the CodecMOS-Accent dataset, a mean opinion score (MOS) benchmark designed to evaluate neural audio codec (NAC) models and the large language model (LLM)-based text-to-s…

eess.AS2025

Good practices for evaluation of synthesized speech

Erica Cooper, Sébastien Le Maguer, Esther Klabbers +1

This document is provided as a guideline for reviewers of papers about speech synthesis. We outline some best practices and common pitfalls for papers about speech synthesis, with…

cs.SD2021

How do Voices from Past Speech Synthesis Challenges Compare Today?

Erica Cooper, Junichi Yamagishi

Shared challenges provide a venue for comparing systems trained on common data using a standardized evaluation, and they also provide an invaluable resource for researchers when th…

eess.AS2024

Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio

Lin Zhang, Xin Wang, Erica Cooper +4

This paper defines Spoof Diarization as a novel task in the Partial Spoof (PS) scenario. It aims to determine what spoofed when, which includes not only locating spoof regions but…

eess.AS2023

The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance

Lin Zhang, Xin Wang, Erica Cooper +2

Automatic speaker verification is susceptible to various manipulations and spoofing, such as text-to-speech synthesis, voice conversion, replay, tampering, adversarial attacks, and…

cs.SD2025

Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores

Jingjing Tang, Erica Cooper, Xin Wang +2

This paper presents an integrated system that transforms symbolic music scores into expressive piano performance audio. By combining a Transformer-based Expressive Performance Rend…

eess.AS2023

Partial Rank Similarity Minimization Method for Quality MOS Prediction of Unseen Speech Synthesis Systems in Zero-Shot and Semi-supervised setting

Hemant Yadav, Erica Cooper, Junichi Yamagishi +2

This paper introduces a novel objective function for quality mean opinion score (MOS) prediction of unseen speech synthesis systems. The proposed function measures the similarity o…

eess.AS2022

Joint Speaker Encoder and Neural Back-end Model for Fully End-to-End Automatic Speaker Verification with Multiple Enrollment Utterances

Chang Zeng, Xiaoxiao Miao, Xin Wang +2

Conventional automatic speaker verification systems can usually be decomposed into a front-end model such as time delay neural network (TDNN) for extracting speaker embeddings and…

eess.AS2020

Improved Prosody from Learned F0 Codebook Representations for VQ-VAE Speech Waveform Reconstruction

Yi Zhao, Haoyu Li, Cheng-I Lai +3

Vector Quantized Variational AutoEncoders (VQ-VAE) are a powerful representation learning framework that can discover discrete groups of features from a speech signal without super…

cs.CL2021

An Investigation of the Relation Between Grapheme Embeddings and Pronunciation for Tacotron-based Systems

Antoine Perquin, Erica Cooper, Junichi Yamagishi

End-to-end models, particularly Tacotron-based ones, are currently a popular solution for text-to-speech synthesis. They allow the production of high-quality synthesized speech wit…

cs.SD2023

SynVox2: Towards a privacy-friendly VoxCeleb2 dataset

Xiaoxiao Miao, Xin Wang, Erica Cooper +5

The success of deep learning in speaker recognition relies heavily on the use of large datasets. However, the data-hungry nature of deep learning methods has already being question…

eess.AS2021

An Initial Investigation for Detecting Partially Spoofed Audio

Lin Zhang, Xin Wang, Erica Cooper +3

All existing databases of spoofed speech contain attack data that is spoofed in its entirety. In practice, it is entirely plausible that successful attacks can be mounted with utte…

cs.SD2025

SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit

Wen-Chin Huang, Erica Cooper, Tomoki Toda

We introduce SHEET, a multi-purpose open-source toolkit designed to accelerate subjective speech quality assessment (SSQA) research. SHEET stands for the Speech Human Evaluation Es…

eess.AS2021

Exploring Disentanglement with Multilingual and Monolingual VQ-VAE

Jennifer Williams, Jason Fong, Erica Cooper +1

This work examines the content and usefulness of disentangled phone and speaker representations from two separately trained VQ-VAE systems: one trained on multilingual data and ano…

eess.AS2023

The VoiceMOS Challenge 2023: Zero-shot Subjective Speech Quality Prediction for Multiple Domains

Erica Cooper, Wen-Chin Huang, Yu Tsao +3

We present the second edition of the VoiceMOS Challenge, a scientific event that aims to promote the study of automatic prediction of the mean opinion score (MOS) of synthesized an…

cs.SD2022

Language-Independent Speaker Anonymization Approach using Self-Supervised Pre-Trained Models

Xiaoxiao Miao, Xin Wang, Erica Cooper +2

Speaker anonymization aims to protect the privacy of speakers while preserving spoken linguistic information from speech. Current mainstream neural network speaker anonymization sy…

eess.AS2025

Layer-wise Analysis for Quality of Multilingual Synthesized Speech

Erica Cooper, Takuma Okamoto, Yamato Ohtani +2

While supervised quality predictors for synthesized speech have demonstrated strong correlations with human ratings, their requirement for in-domain labeled training data hinders t…

eess.AS2021

Attention Back-end for Automatic Speaker Verification with Multiple Enrollment Utterances

Chang Zeng, Xin Wang, Erica Cooper +2

Probabilistic linear discriminant analysis (PLDA) or cosine similarity have been widely used in traditional speaker verification systems as back-end techniques to measure pairwise…

cs.SD2023

Exploring Isolated Musical Notes as Pre-training Data for Predominant Instrument Recognition in Polyphonic Music

Lifan Zhong, Erica Cooper, Junichi Yamagishi +1

With the growing amount of musical data available, automatic instrument recognition, one of the essential problems in Music Information Retrieval (MIR), is drawing more and more at…