6 papers
Brick-Composer: Using MLLMs for Assembly with Diverse Bricks
Jiateng Liu, Bingxuan Li, Zhenhailong Wang +8
We dream of AI agents that can read arbitrary designs and construct real-world objects from reusable building blocks. As a first step toward this vision, we study whether multimoda…
DiscreteRTC: Discrete Diffusion Policies are Natural Asynchronous Executors
Pengcheng Wang, Kaiwen Hong, Chensheng Peng +4
Unlike chatbots, physical AI must act while the world keeps evolving. Therefore, the inter-chunk pause of synchronous executors are fatal for dynamic tasks regardless of how fast t…
Multi-Modal Manipulation via Multi-Modal Policy Consensus
Haonan Chen, Jiaming Xu, Hongyu Chen +7
Effectively integrating diverse sensory modalities is crucial for robotic manipulation. However, the typical approach of feature concatenation is often suboptimal: dominant modalit…
HEIGHT: Heterogeneous Interaction Graph Transformer for Robot Navigation in Crowded and Constrained Environments
Shuijing Liu, Haochen Xia, Fatemeh Cheraghi Pouria +5
We study the problem of robot navigation in dense and interactive crowds with static constraints such as corridors and furniture. Previous methods fail to consider all types of spa…
Human-Agent Joint Learning for Efficient Robot Manipulation Skill Acquisition
Shengcheng Luo, Quanquan Peng, Jun Lv +4
Employing a teleoperation system for gathering demonstrations offers the potential for more efficient learning of robot manipulation. However, teleoperating a robot arm equipped wi…
Structured Graph Network for Constrained Robot Crowd Navigation with Low Fidelity Simulation
Shuijing Liu, Kaiwen Hong, Neeloy Chakraborty +1
We investigate the feasibility of deploying reinforcement learning (RL) policies for constrained crowd navigation using a low-fidelity simulator. We introduce a representation of t…