activity
20192026
most citedImproved Prosody from Learned F0 Codebook Representations for VQ-VAE Speech Waveform Reconstruction

3 citations · 5 across the 7 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2026

Disentangling Speaker Traits for Deepfake Source Verification via Chebyshev Polynomial and Riemannian Metric Learning

Xi Xuan, Wenxin Zhang, Zhiyu Li +3

Speech deepfake source verification systems aims to determine whether two synthetic speech utterances originate from the same source generator, often assuming that the resulting so…

eess.AS20221 cited

Analysis of Voice Conversion and Code-Switching Synthesis Using VQ-VAE

Shuvayanti Das, Jennifer Williams, Catherine Lai

This paper presents an analysis of speech synthesis quality achieved by simultaneously performing voice conversion and language code-switching using multilingual VQ-VAE speech synt…

eess.AS2021

Revisiting Speech Content Privacy

Jennifer Williams, Junichi Yamagishi, Paul-Gauthier Noe +2

In this paper, we discuss an important aspect of speech privacy: protecting spoken content. New capabilities from the field of machine learning provide a unique and timely opportun…

eess.AS20211 cited

Exploring Disentanglement with Multilingual and Monolingual VQ-VAE

Jennifer Williams, Jason Fong, Erica Cooper +1

This work examines the content and usefulness of disentangled phone and speaker representations from two separately trained VQ-VAE systems: one trained on multilingual data and ano…

eess.AS2020

Learning Disentangled Phone and Speaker Representations in a Semi-Supervised VQ-VAE Paradigm

Jennifer Williams, Yi Zhao, Erica Cooper +1

We present a new approach to disentangle speaker voice and phone content by introducing new components to the VQ-VAE architecture for speech synthesis. The original VQ-VAE does not…

eess.AS20203 cited

Improved Prosody from Learned F0 Codebook Representations for VQ-VAE Speech Waveform Reconstruction

Yi Zhao, Haoyu Li, Cheng-I Lai +3

Vector Quantized Variational AutoEncoders (VQ-VAE) are a powerful representation learning framework that can discover discrete groups of features from a speech signal without super…