1 citations · 1 across the 3 of their papers we have counts for
3 papers
eess.AS2025
Switchboard-Affect: Emotion Perception Labels from Conversational Speech
Amrit Romana, Jaya Narain, Tien Dung Tran +4
Understanding the nuances of speech emotion dataset curation and labeling is essential for assessing speech emotion recognition (SER) model potential in real-world applications. Mo…
cs.SD2022
Towards zero-shot Text-based voice editing using acoustic context conditioning, utterance embeddings, and reference encoders
Jason Fong, Yun Wang, Prabhav Agrawal +4
Text-based voice editing (TBVE) uses synthetic output from text-to-speech (TTS) systems to replace words in an original recording. Recent work has used neural models to produce edi…
eess.AS2021★ 1 cited
Exploring Disentanglement with Multilingual and Monolingual VQ-VAE
Jennifer Williams, Jason Fong, Erica Cooper +1
This work examines the content and usefulness of disentangled phone and speaker representations from two separately trained VQ-VAE systems: one trained on multilingual data and ano…