activity
20192023
most citedLDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech

10 citations · 29 across the 12 of their papers we have counts for

collaborators

17 papers

cs.SD2023

Exploring Isolated Musical Notes as Pre-training Data for Predominant Instrument Recognition in Polyphonic Music

Lifan Zhong, Erica Cooper, Junichi Yamagishi +1

With the growing amount of musical data available, automatic instrument recognition, one of the essential problems in Music Information Retrieval (MIR), is drawing more and more at…

cs.SD2022

Analyzing Language-Independent Speaker Anonymization Framework under Unseen Conditions

Xiaoxiao Miao, Xin Wang, Erica Cooper +2

In our previous work, we proposed a language-independent speaker anonymization system based on self-supervised learning models. Although the system can anonymize speech data of any…

cs.SD2022

Language-Independent Speaker Anonymization Approach using Self-Supervised Pre-Trained Models

Xiaoxiao Miao, Xin Wang, Erica Cooper +2

Speaker anonymization aims to protect the privacy of speakers while preserving spoken linguistic information from speech. Current mainstream neural network speaker anonymization sy…

cs.SD2021

On the Interplay Between Sparsity, Naturalness, Intelligibility, and Prosody in Speech Synthesis

Cheng-I Jeff Lai, Erica Cooper, Yang Zhang +8

Are end-to-end text-to-speech (TTS) models over-parametrized? To what extent can these models be pruned, and what happens to their synthesis capabilities? This work serves as a sta…

cs.SD202110 cited

LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech

Wen-Chin Huang, Erica Cooper, Junichi Yamagishi +1

An effective approach to automatically predict the subjective rating for synthetic speech is to train on a listening test dataset with human-annotated scores. Although each speech…

cs.SD20211 cited

Multi-Task Learning in Utterance-Level and Segmental-Level Spoof Detection

Lin Zhang, Xin Wang, Erica Cooper +1

In this paper, we provide a series of multi-tasking benchmarks for simultaneously detecting spoofing at the segmental and utterance levels in the PartialSpoof database. First, we p…