13 papers
Tired Actor: Fatigue-Informed Character Control
Shengyuan Zhang, Xinpeng Liu, Muchun Niu +5
Replicating human behavior with physics simulation has been a long-expected goal in character animation. Existing efforts have achieved impressive performance in imitating a wide s…
Explicit Kinematic Guidance from Analytic Concepts for Vision-Language-Action Models
Mingyang Sun, Jiude Wei, Xiujian Liang +4
The paper introduces a Concept Expert module that extracts 3D kinematic and structural information to create Analytic Concepts, providing explicit guidance for Vision-Language-Acti…
Analytic Concept-Centric Memory for Agentic Embodied Manipulation
Mingyang Sun, Xiujian Liang, Jiude Wei +4
Long-horizon embodied manipulation requires agents to remember persistent objects, track changing scene states, and reuse prior interaction knowledge. However, existing agent memor…
Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts
Jiude Wei, Yuxuan Li, Cewu Lu +1
We humans rely on a wide range of commonsense knowledge to interact with an extensive number and categories of objects in the physical world. Likewise, such commonsense knowledge i…
Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models
Zehao Wang, Xinpeng Liu, Yudonglin Zhang +6
Multimodal Large Language Models (MLLMs) have garnered significant attention recently and demonstrate outstanding capabilities in various tasks such as OCR, VQA, captioning, $\text…
RealDiff: Bridging Real-World Gap in Robot Manipulation via Depth Diffusion
Xiujian Liang, Jiacheng Liu, Mingyang Sun +3
Robot manipulation in the real world is fundamentally constrained by the visual sim2real gap, where depth observations collected in simulation fail to reflect the complex noise pat…