4 citations · 4 across the 2 of their papers we have counts for
12 papers
Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis
Yu Qi, Haibo Zhao, Ziyu Guo +17
Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, existing embodied benchmarks mainly…
ThinkGrasp: A Vision-Language System for Strategic Part Grasping in Clutter
Yaoyao Qian, Xupeng Zhu, Ondrej Biza +5
Robotic grasping in cluttered environments remains a significant challenge due to occlusions and complex object arrangements. We have developed ThinkGrasp, a plug-and-play vision-l…
3D Equivariant Visuomotor Policy Learning via Spherical Projection
Boce Hu, Dian Wang, David Klee +5
Equivariant models have recently been shown to improve the data efficiency of diffusion policy by a significant margin. However, prior work that explored this direction focused pri…
Generalizable Hierarchical Skill Learning via Object-Centric Representation
Haibo Zhao, Yu Qi, Boce Hu +9
We present Generalizable Hierarchical Skill Learning (GSL), a novel framework for hierarchical policy learning that significantly improves policy generalization and sample efficien…
Clebsch-Gordan Transformer: Fast and Global Equivariant Attention
Owen Lewis Howell, Linfeng Zhao, Xupeng Zhu +6
The global attention mechanism is one of the keys to the success of transformer architecture, but it incurs quadratic computational costs in relation to the number of tokens. On th…
SE(3)-Equivariant Diffusion Policy in Spherical Fourier Space
Xupeng Zhu, Fan Wang, Robin Walters +1
Diffusion Policies are effective at learning closed-loop manipulation policies from human demonstrations but generalize poorly to novel arrangements of objects in 3D space, hurting…