collaborators

6 papers

cs.CV2026

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

Yu Qi, Haibo Zhao, Ziyu Guo +17

Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, existing embodied benchmarks mainly…

cs.CV2026

SMAL-pets: SMAL Based Avatars of Pets from Single Image

Piotr Borycki, Joanna Waczyńska, Yizhe Zhu +2

Creating high-fidelity, animatable 3D dog avatars remains a formidable challenge in computer vision. Unlike human digital doubles, animal reconstruction faces a critical shortage o…

cs.RO2025

Residual Rotation Correction using Tactile Equivariance

Yizhe Zhu, Zhang Ye, Boce Hu +4

Visuotactile policy learning augments vision-only policies with tactile input, facilitating contact-rich manipulation. However, the high cost of tactile data collection makes sampl…

cs.RO2025

Generalizable Hierarchical Skill Learning via Object-Centric Representation

Haibo Zhao, Yu Qi, Boce Hu +9

We present Generalizable Hierarchical Skill Learning (GSL), a novel framework for hierarchical policy learning that significantly improves policy generalization and sample efficien…

cs.RO2025

EquAct: An SE(3)-Equivariant Multi-Task Transformer for Open-Loop Robotic Manipulation

Xupeng Zhu, Yu Qi, Yizhe Zhu +2

Transformer architectures can effectively learn language-conditioned, multi-task 3D open-loop manipulation policies from demonstrations by jointly processing natural language instr…

cs.RO2025

Hierarchical Equivariant Policy via Frame Transfer

Haibo Zhao, Dian Wang, Yizhe Zhu +6

Recent advances in hierarchical policy learning highlight the advantages of decomposing systems into high-level and low-level agents, enabling efficient long-horizon reasoning and…