15 papers
DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
Yikai Tang, Haoran Geng, Jindou Jia +5
Imitation learning has emerged as a crucial approach for acquiring visuomotor skills from demonstrations, where designing effective observation encoders is essential for policy gen…
Multi-Objective Learning for Diffusion Models: A Statistical Theory under Semi-Supervised Learning
Ziheng Cheng, Yixiao Huang, Hanlin Zhu +5
Diffusion models are increasingly used as powerful conditional generators, yet real deployments often involve multiple target distributions arising from different tasks, e.g., dive…
ViTacFormer: Learning Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation
Liang Heng, Haoran Geng, Kaifeng Zhang +2
Dexterous manipulation is a cornerstone capability for robotic systems aiming to interact with the physical world in a human-like manner. Although vision-based methods have advance…
Large Video Planner Enables Generalizable Robot Control
Boyuan Chen, Tianyuan Zhang, Haoran Geng +9
General-purpose robots require decision-making models that generalize across diverse tasks and environments. Recent works build robot foundation models by extending multimodal larg…
World Model for Robot Learning: A Comprehensive Survey
Bohan Hou, Gen Li, Jindou Jia +15
World models, which are predictive representations of how environments evolve under actions, have become a central component of robot learning. They support policy learning, planni…
Rodrigues Network for Learning Robot Actions
Jialiang Zhang, Haoran Geng, Yang You +4
Understanding and predicting articulated actions is important in robot learning. However, common architectures such as MLPs and Transformers lack inductive biases that reflect the…