4 papers
Diffusion Models are Open-World Affordance Learners: Leveraging Generative Priors for 3D Affordance Learning
Hanqing Wang, Zhenhao Zhang, Kaiyang Ji +12
3D affordance grounding aims to understand how diverse objects can be manipulated, making it a cornerstone of embodied interaction. However, prior works struggle to generalize to o…
From Scene-Centric to Observer-Centric: Modeling Observer-Aware Relations for 3D Scene Graph Generation
Jingjun Sun, Chaowei Wang, Zhirui Liu +5
3D Scene Graph Generation (3DSGG) represents 3D scenes as structured object--relation--object graphs for spatial understanding. In observer-centric spatial perception, the same sce…
Commanding Humanoid by Free-form Language: A Large Language Action Model with Unified Motion Vocabulary
Zhirui Liu, Kaiyang Ji, Ke Yang +4
Enabling humanoid robots to follow free-form natural language commands is a critical step toward seamless human-robot interaction and general-purpose embodied AI. However, existing…
HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model
Zhenhao Zhang, Hanqing Wang, Xiangyu Zeng +10
Understanding and recognizing human-object interaction (HOI) is a pivotal application in AR/VR and robotics. Recent open-vocabulary HOI detection approaches depend exclusively on l…