activity
20112024
most citedAUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss

195 citations · 446 across the 17 of their papers we have counts for

collaborators

30 papers

eess.AS20221 cited

WAVPROMPT: Towards Few-Shot Spoken Language Understanding with Frozen Language Models

Heting Gao, Junrui Ni, Kaizhi Qian +3

Large-scale auto-regressive language models pretrained on massive text have demonstrated their impressive ability to perform new natural language tasks with only a few text example…

cs.LG20223 cited

Equivariance Discovery by Learned Parameter-Sharing

Raymond A. Yeh, Yuan-Ting Hu, Mark Hasegawa-Johnson +1

Designing equivariance as an inductive bias into deep-nets has been a prominent approach to build effective models, e.g., a convolutional neural network incorporates translation eq…

eess.AS2022

Visualizations of Complex Sequences of Family-Infant Vocalizations Using Bag-of-Audio-Words Approach Based on Wav2vec 2.0 Features

Jialu Li, Mark Hasegawa-Johnson, Nancy L. McElwain

In the U.S., approximately 15-17% of children 2-8 years of age are estimated to have at least one diagnosed mental, behavioral or developmental disorder. However, such disorders of…

eess.AS2022

SpeechSplit 2.0: Unsupervised speech disentanglement for voice conversion Without tuning autoencoder Bottlenecks

Chak Ho Chan, Kaizhi Qian, Yang Zhang +1

SpeechSplit can perform aspect-specific voice conversion by disentangling speech into content, rhythm, pitch, and timbre using multiple autoencoders in an unsupervised manner. Howe…

cs.SD2022

Discovering Phonetic Inventories with Crosslingual Automatic Speech Recognition

Piotr Żelasko, Siyuan Feng, Laureano Moro Velazquez +5

The high cost of data acquisition makes Automatic Speech Recognition (ASR) model training problematic for most existing languages, including languages that do not even have a writt…

eess.AS202112 cited

Global Rhythm Style Transfer Without Text Transcriptions

Kaizhi Qian, Yang Zhang, Shiyu Chang +4

Prosody plays an important role in characterizing the style of a speaker or an emotion, but most non-parallel voice or emotion style transfer algorithms do not convert any prosody…