activity
20192022
most citedImproved Prosody from Learned F0 Codebook Representations for VQ-VAE Speech Waveform Reconstruction

3 citations · 5 across the 5 of their papers we have counts for

collaborators

8 papers

cs.CR2022

Attacker Attribution of Audio Deepfakes

Nicolas M. Müller, Franziska Dieckmann, Jennifer Williams

Deepfakes are synthetically generated media often devised with malicious intent. They have become increasingly more convincing with large training datasets advanced neural networks…

eess.AS20221 cited

Analysis of Voice Conversion and Code-Switching Synthesis Using VQ-VAE

Shuvayanti Das, Jennifer Williams, Catherine Lai

This paper presents an analysis of speech synthesis quality achieved by simultaneously performing voice conversion and language code-switching using multilingual VQ-VAE speech synt…

eess.AS2021

Revisiting Speech Content Privacy

Jennifer Williams, Junichi Yamagishi, Paul-Gauthier Noe +2

In this paper, we discuss an important aspect of speech privacy: protecting spoken content. New capabilities from the field of machine learning provide a unique and timely opportun…

eess.AS20211 cited

Exploring Disentanglement with Multilingual and Monolingual VQ-VAE

Jennifer Williams, Jason Fong, Erica Cooper +1

This work examines the content and usefulness of disentangled phone and speaker representations from two separately trained VQ-VAE systems: one trained on multilingual data and ano…

eess.AS2020

Learning Disentangled Phone and Speaker Representations in a Semi-Supervised VQ-VAE Paradigm

Jennifer Williams, Yi Zhao, Erica Cooper +1

We present a new approach to disentangle speaker voice and phone content by introducing new components to the VQ-VAE architecture for speech synthesis. The original VQ-VAE does not…

eess.AS20203 cited

Improved Prosody from Learned F0 Codebook Representations for VQ-VAE Speech Waveform Reconstruction

Yi Zhao, Haoyu Li, Cheng-I Lai +3

Vector Quantized Variational AutoEncoders (VQ-VAE) are a powerful representation learning framework that can discover discrete groups of features from a speech signal without super…