2 papers
cs.HC2024
A Unified Editing Method for Co-Speech Gesture Generation via Diffusion Inversion
Zeyu Zhao, Nan Gao, Zhi Zeng +3
Diffusion models have shown great success in generating high-quality co-speech gestures for interactive humanoid robots or digital avatars from noisy input with the speech audio or…
eess.AS2023
ASR and Emotional Speech: A Word-Level Investigation of the Mutual Impact of Speech and Emotion Recognition
Yuanchao Li, Zeyu Zhao, Ondrej Klejch +2
In Speech Emotion Recognition (SER), textual data is often used alongside audio signals to address their inherent variability. However, the reliance on human annotated text in most…