2 papers
cs.GR2024
LLM Gesticulator: Leveraging Large Language Models for Scalable and Controllable Co-Speech Gesture Synthesis
Haozhou Pang, Tianwei Ding, Lanshan He +3
In this work, we present LLM Gesticulator, an LLM-based audio-driven co-speech gesture generation framework that synthesizes full-body animations that are rhythmically aligned with…
cs.CV2024
Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout
Anbin QI, Zhongliang Liu, Xinyong Zhou +6
In this paper, we present our solution for the Second Multimodal Emotion Recognition Challenge Track 1(MER2024-SEMI). To enhance the accuracy and generalization performance of emot…