From the 1 of 7 linked papers with an AI index.
7 papers
Retriever: Composing Closed-Loop Asynchronous Robot Programs
Linfeng Zhao, Haojie Huang, Jiayuan Mao +3
Building long-horizon robot agents requires composing closed-loop pipelines -- perception, belief update, planning, and control -- whose components run at different clocks and with…
Pix2Act: Image-Space Manipulation Policies with Equivariant Augmentation
Haojie Huang, Linfeng Zhao, Haotian Liu +9
Pix2Act is an imitation‑learning approach that predicts continuous 2D keypoint trajectories in camera images and recovers 3D end‑effector poses via triangulation, using equivariant…
Action Map Policy: Learning 3D Closed-loop Manipulation via Pixel Classification
Haojie Huang, Zhang Ye, Linfeng Zhao +7
The action space poses a major challenge in robot learning, since it is often high-dimensional, can span long time horizons, and frequently admits multi-modal optimal solutions. A…
ThinkGrasp: A Vision-Language System for Strategic Part Grasping in Clutter
Yaoyao Qian, Xupeng Zhu, Ondrej Biza +5
Robotic grasping in cluttered environments remains a significant challenge due to occlusions and complex object arrangements. We have developed ThinkGrasp, a plug-and-play vision-l…
Clebsch-Gordan Transformer: Fast and Global Equivariant Attention
Owen Lewis Howell, Linfeng Zhao, Xupeng Zhu +6
The global attention mechanism is one of the keys to the success of transformer architecture, but it incurs quadratic computational costs in relation to the number of tokens. On th…
Learning Efficient and Robust Language-conditioned Manipulation using Textual-Visual Relevancy and Equivariant Language Mapping
Mingxi Jia, Haojie Huang, Zhewen Zhang +7
Controlling robots through natural language is pivotal for enhancing human-robot collaboration and synthesizing complex robot behaviors. Recent works that are trained on large robo…