From the 1 of 16 linked papers with an AI index.
16 papers
Pix2Act: Image-Space Manipulation Policies with Equivariant Augmentation
Haojie Huang, Linfeng Zhao, Haotian Liu +9
Pix2Act is an imitation‑learning approach that predicts continuous 2D keypoint trajectories in camera images and recovers 3D end‑effector poses via triangulation, using equivariant…
Action Map Policy: Learning 3D Closed-loop Manipulation via Pixel Classification
Haojie Huang, Zhang Ye, Linfeng Zhao +7
The action space poses a major challenge in robot learning, since it is often high-dimensional, can span long time horizons, and frequently admits multi-modal optimal solutions. A…
ThinkGrasp: A Vision-Language System for Strategic Part Grasping in Clutter
Yaoyao Qian, Xupeng Zhu, Ondrej Biza +5
Robotic grasping in cluttered environments remains a significant challenge due to occlusions and complex object arrangements. We have developed ThinkGrasp, a plug-and-play vision-l…
A Practical Guide for Incorporating Symmetry in Diffusion Policy
Dian Wang, Boce Hu, Shuran Song +2
Recently, equivariant neural networks for policy learning have shown promising improvements in sample efficiency and generalization, however, their wide adoption faces substantial…
Residual Rotation Correction using Tactile Equivariance
Yizhe Zhu, Zhang Ye, Boce Hu +4
Visuotactile policy learning augments vision-only policies with tactile input, facilitating contact-rich manipulation. However, the high cost of tactile data collection makes sampl…
3D Equivariant Visuomotor Policy Learning via Spherical Projection
Boce Hu, Dian Wang, David Klee +5
Equivariant models have recently been shown to improve the data efficiency of diffusion policy by a significant margin. However, prior work that explored this direction focused pri…