activity
20242026
collaborators

5 papers

cs.GR2026

Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation

Lanshan He, Haozhou Pang, Qi Gan +12

Cutscenes are carefully choreographed cinematic sequences embedded in video games and interactive media, serving as the primary vehicle for narrative delivery, character developmen…

cs.CV2026

OmniDiT: Extending Diffusion Transformer to Omni-VTON Framework

Weixuan Zeng, Pengcheng Wei, Huaiqing Wang +8

Despite the rapid advancement of Virtual Try-On (VTON) and Try-Off (VTOFF) technologies, existing VTON methods face challenges with fine-grained detail preservation, generalization…

cs.GR2025

Global Position Aware Group Choreography using Large Language Model

Haozhou Pang, Tianwei Ding, Lanshan He +1

Dance serves as a profound and universal expression of human culture, conveying emotions and stories through movements synchronized with music. Although some current works have ach…

cs.GR2024

LLM Gesticulator: Leveraging Large Language Models for Scalable and Controllable Co-Speech Gesture Synthesis

Haozhou Pang, Tianwei Ding, Lanshan He +3

In this work, we present LLM Gesticulator, an LLM-based audio-driven co-speech gesture generation framework that synthesizes full-body animations that are rhythmically aligned with…

cs.CV2024

Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout

Anbin QI, Zhongliang Liu, Xinyong Zhou +6

In this paper, we present our solution for the Second Multimodal Emotion Recognition Challenge Track 1(MER2024-SEMI). To enhance the accuracy and generalization performance of emot…