11 papers
KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation
Xinyu Shao, Keru Zhou, Guowei Huang +3
Learning manipulation from few demonstrations requires visual priors that capture not only where to interact, but also how the interaction should begin; static priors such as segme…
Whole-Body Inverse Kinematics with Graph Diffusion
Helong Huang, Kai Tan, Feng Wen +2
Inverse kinematics (IK) is a fundamental problem in robotics, requiring the generation of joint configurations that satisfy target end-effector poses. Existing approaches often str…
ForceFlow: Learning to Feel and Act via Contact-Driven Flow Matching
Shuoheng Zhang, Yifu Yuan, Hongyao Tang +7
Existing imitation learning methods enable robots to interact autonomously with the physical environment. However, contact-rich manipulation tasks remain a significant challenge du…
UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
Yueming Xu, Jiahui Zhang, Ze Huang +12
Despite the impressive progress on understanding and generating images shown by the recent unified architectures, the integration of 3D tasks remains challenging and largely unexpl…
RADAR: Revealing Asymmetric Development of Abilities in MLLM Pre-training
Yunshuang Nie, Bingqian Lin, Minzhe Niu +7
Pre-trained Multi-modal Large Language Models (MLLMs) provide a knowledge-rich foundation for post-training by leveraging their inherent perception and reasoning capabilities to so…
From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D
Jiahui Zhang, Yurui Chen, Yanpeng Zhou +10
Recent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unl…