195 citations · 446 across the 13 of their papers we have counts for
4 papers · 1 filter
WAVPROMPT: Towards Few-Shot Spoken Language Understanding with Frozen Language Models
Heting Gao, Junrui Ni, Kaizhi Qian +3
Large-scale auto-regressive language models pretrained on massive text have demonstrated their impressive ability to perform new natural language tasks with only a few text example…
Global Rhythm Style Transfer Without Text Transcriptions
Kaizhi Qian, Yang Zhang, Shiyu Chang +4
Prosody plays an important role in characterizing the style of a speaker or an emotion, but most non-parallel voice or emotion style transfer algorithms do not convert any prosody…
Unsupervised Speech Decomposition via Triple Information Bottleneck
Kaizhi Qian, Yang Zhang, Shiyu Chang +2
Speech information can be roughly decomposed into four components: language content, timbre, pitch, and rhythm. Obtaining disentangled representations of these components is useful…
AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
Kaizhi Qian, Yang Zhang, Shiyu Chang +2
Non-parallel many-to-many voice conversion, as well as zero-shot voice conversion, remain under-explored areas. Deep style transfer algorithms, such as generative adversarial netwo…