works on

From the 1 of 11 linked papers with an AI index.

collaborators

11 papers

cs.RO2026

Pix2Act: Image-Space Manipulation Policies with Equivariant Augmentation

Haojie Huang, Linfeng Zhao, Haotian Liu +9

Pix2Act is an imitation‑learning approach that predicts continuous 2D keypoint trajectories in camera images and recovers 3D end‑effector poses via triangulation, using equivariant…

cs.RO2026

Action Map Policy: Learning 3D Closed-loop Manipulation via Pixel Classification

Haojie Huang, Zhang Ye, Linfeng Zhao +7

The action space poses a major challenge in robot learning, since it is often high-dimensional, can span long time horizons, and frequently admits multi-modal optimal solutions. A…

cs.CV2026

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

Yu Qi, Haibo Zhao, Ziyu Guo +17

Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, existing embodied benchmarks mainly…

cs.RO2026

ThinkGrasp: A Vision-Language System for Strategic Part Grasping in Clutter

Yaoyao Qian, Xupeng Zhu, Ondrej Biza +5

Robotic grasping in cluttered environments remains a significant challenge due to occlusions and complex object arrangements. We have developed ThinkGrasp, a plug-and-play vision-l…

cs.RO2025

3D Equivariant Visuomotor Policy Learning via Spherical Projection

Boce Hu, Dian Wang, David Klee +5

Equivariant models have recently been shown to improve the data efficiency of diffusion policy by a significant margin. However, prior work that explored this direction focused pri…

cs.RO2025

Generalizable Hierarchical Skill Learning via Object-Centric Representation

Haibo Zhao, Yu Qi, Boce Hu +9

We present Generalizable Hierarchical Skill Learning (GSL), a novel framework for hierarchical policy learning that significantly improves policy generalization and sample efficien…