17 papers
Leverage Cross-Attention for End-to-End Open-Vocabulary Panoptic Reconstruction
Xuan Yu, Yuxuan Xie, Yili Liu +4
Open-vocabulary panoptic reconstruction offers comprehensive scene understanding, enabling advances in embodied robotics and photorealistic simulation. In this paper, we propose Pa…
UnIRe: Unsupervised Instance Decomposition for Dynamic Urban Scene Reconstruction
Yunxuan Mao, Rong Xiong, Yue Wang +1
Reconstructing and decomposing dynamic urban scenes is crucial for autonomous driving, urban planning, and scene editing. However, existing methods fail to perform instance-aware d…
Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter
Kechun Xu, Xunlong Xia, Kaixuan Wang +6
We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches le…
Disambiguate Gripper State in Grasp-Based Tasks: Pseudo-Tactile as Feedback Enables Pure Simulation Learning
Yifei Yang, Lu Chen, Zherui Song +5
Grasp-based manipulation tasks are fundamental to robots interacting with their environments, yet gripper state ambiguity significantly reduces the robustness of imitation learning…
PanopticSplatting: End-to-End Panoptic Gaussian Splatting
Yuxuan Xie, Xuan Yu, Changjian Jiang +5
Open-vocabulary panoptic reconstruction is a challenging task for simultaneous scene reconstruction and understanding. Recently, methods have been proposed for 3D scene understandi…
Natural Humanoid Robot Locomotion with Generative Motion Prior
Haodong Zhang, Liang Zhang, Zhenghan Chen +3
Natural and lifelike locomotion remains a fundamental challenge for humanoid robots to interact with human society. However, previous methods either neglect motion naturalness or r…