6 papers
Bridging Modality Disconnect in Self-Reflection via Closed-Loop Visually Grounded Verification
Haoyu Zhang, Yuwei Wu, Pengxiang Li +6
In the era of Vision-Language Models (VLMs), enhancing multimodal reasoning capabilities remains a critical challenge, particularly in handling ambiguous or complex visual inputs,…
Facial Expression Generation Aligned with Human Preference for Natural Dyadic Interaction
Xu Chen, Rui Gao, Xinjie Zhang +5
Achieving natural dyadic interaction requires generating facial expressions that are emotionally appropriate and socially aligned with human preference. Human feedback offers a com…
Morphology-Independent Facial Expression Imitation for Human-Face Robots
Xu Chen, Rui Gao, Che Sun +4
Accurate facial expression imitation on human-face robots is crucial for achieving natural human-robot interaction. Most existing methods have achieved photorealistic expression im…
Fine-Grained 3D Facial Reconstruction for Micro-Expressions
Che Sun, Xinjie Zhang, Rui Gao +3
Recent advances in 3D facial expression reconstruction have demonstrated remarkable performance in capturing macro-expressions, yet the reconstruction of micro-expressions remains…
Long-Horizon Visual Imitation Learning via Plan and Code Reflection
Quan Chen, Chenrui Shi, Qi Chen +6
Learning from long-horizon demonstrations with complex action sequences presents significant challenges for visual imitation learning, particularly in understanding temporal relati…
M3PO: Multimodal-Model-Guided Preference Optimization for Visual Instruction Following
Ruirui Gao, Emily Johnson, Bowen Tan +1
Large Vision-Language Models (LVLMs) hold immense potential for complex multimodal instruction following, yet their development is often hindered by the high cost and inconsistency…