From the 2 of 13 linked papers with an AI index.
13 papers
Mixture of Frames Policy: Multi-Frame Action Denoising for Bimanual Mobile Manipulation
Dian Wang, Jisang Park, Xiaomeng Xu +3
The paper introduces Mixture of Frames Policy (MoF), a diffusion-based visuomotor controller that denoises actions simultaneously in several coordinate frames to better handle bima…
Pix2Act: Image-Space Manipulation Policies with Equivariant Augmentation
Haojie Huang, Linfeng Zhao, Haotian Liu +9
Pix2Act is an imitation‑learning approach that predicts continuous 2D keypoint trajectories in camera images and recovers 3D end‑effector poses via triangulation, using equivariant…
Action Map Policy: Learning 3D Closed-loop Manipulation via Pixel Classification
Haojie Huang, Zhang Ye, Linfeng Zhao +7
The action space poses a major challenge in robot learning, since it is often high-dimensional, can span long time horizons, and frequently admits multi-modal optimal solutions. A…
HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations
Xiaomeng Xu, Jisang Park, Han Zhang +6
We present Whole-Body Mobile Manipulation Interface (HoMMI), a data collection and policy learning framework that learns whole-body mobile manipulation directly from robot-free hum…
A Practical Guide for Incorporating Symmetry in Diffusion Policy
Dian Wang, Boce Hu, Shuran Song +2
Recently, equivariant neural networks for policy learning have shown promising improvements in sample efficiency and generalization, however, their wide adoption faces substantial…
3D Equivariant Visuomotor Policy Learning via Spherical Projection
Boce Hu, Dian Wang, David Klee +5
Equivariant models have recently been shown to improve the data efficiency of diffusion policy by a significant margin. However, prior work that explored this direction focused pri…