4 papers
Event-Driven Proactive Assistive Manipulation with Grounded Vision-Language Planning
Fengkai Liu, Hao Su, Haozhuang Chi +6
Assistance in collaborative manipulation is often initiated by user instructions, making high-level reasoning request-driven. In fluent human teamwork, however, partners often infe…
AcoustEmo: Open-Vocabulary Emotion Reasoning via Utterance-Aware Acoustic Q-Former
Liyun Zhang, Xuanmeng Sha, Shuqiong Wu +1
Multimodal Large Language Models (MLLMs) excel in Open-Vocabulary (OV) emotion recognition but often neglect fine-grained acoustic modeling. Existing methods typically use global a…
3DGesPolicy: Phoneme-Aware Holistic Co-Speech Gesture Generation Based on Action Control
Xuanmeng Sha, Liyun Zhang, Tomohiro Mashita +2
Generating holistic co-speech gestures that integrate full-body motion with facial expressions suffers from semantically incoherent coordination on body motion and spatially unstab…
3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control
Xuanmeng Sha, Liyun Zhang, Tomohiro Mashita +2
Audio-driven 3D facial animation has achieved significant progress in both research and applications. While recent baselines struggle to generate natural and continuous facial move…