From the 1 of 11 linked papers with an AI index.
11 papers
Pix2Act: Image-Space Manipulation Policies with Equivariant Augmentation
Haojie Huang, Linfeng Zhao, Haotian Liu +9
Pix2Act is an imitation‑learning approach that predicts continuous 2D keypoint trajectories in camera images and recovers 3D end‑effector poses via triangulation, using equivariant…
Action Map Policy: Learning 3D Closed-loop Manipulation via Pixel Classification
Haojie Huang, Zhang Ye, Linfeng Zhao +7
The action space poses a major challenge in robot learning, since it is often high-dimensional, can span long time horizons, and frequently admits multi-modal optimal solutions. A…
Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis
Yu Qi, Haibo Zhao, Ziyu Guo +17
Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, existing embodied benchmarks mainly…
ThinkGrasp: A Vision-Language System for Strategic Part Grasping in Clutter
Yaoyao Qian, Xupeng Zhu, Ondrej Biza +5
Robotic grasping in cluttered environments remains a significant challenge due to occlusions and complex object arrangements. We have developed ThinkGrasp, a plug-and-play vision-l…
3D Equivariant Visuomotor Policy Learning via Spherical Projection
Boce Hu, Dian Wang, David Klee +5
Equivariant models have recently been shown to improve the data efficiency of diffusion policy by a significant margin. However, prior work that explored this direction focused pri…
Generalizable Hierarchical Skill Learning via Object-Centric Representation
Haibo Zhao, Yu Qi, Boce Hu +9
We present Generalizable Hierarchical Skill Learning (GSL), a novel framework for hierarchical policy learning that significantly improves policy generalization and sample efficien…