1 citations · 1 across the 4 of their papers we have counts for
17 papers
DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
Yikai Tang, Haoran Geng, Jindou Jia +5
Imitation learning has emerged as a crucial approach for acquiring visuomotor skills from demonstrations, where designing effective observation encoders is essential for policy gen…
Multi-Objective Learning for Diffusion Models: A Statistical Theory under Semi-Supervised Learning
Ziheng Cheng, Yixiao Huang, Hanlin Zhu +5
Diffusion models are increasingly used as powerful conditional generators, yet real deployments often involve multiple target distributions arising from different tasks, e.g., dive…
ViTacFormer: Learning Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation
Liang Heng, Haoran Geng, Kaifeng Zhang +2
Dexterous manipulation is a cornerstone capability for robotic systems aiming to interact with the physical world in a human-like manner. Although vision-based methods have advance…
Large Video Planner Enables Generalizable Robot Control
Boyuan Chen, Tianyuan Zhang, Haoran Geng +9
General-purpose robots require decision-making models that generalize across diverse tasks and environments. Recent works build robot foundation models by extending multimodal larg…
Rodrigues Network for Learning Robot Actions
Jialiang Zhang, Haoran Geng, Yang You +4
Understanding and predicting articulated actions is important in robot learning. However, common architectures such as MLPs and Transformers lack inductive biases that reflect the…
OSGym: Scalable OS Infra for Computer Use Agents
Zengyi Qin, Jinyuan Chen, Yunze Man +25
Training computer use agents requires full-featured OS sandboxes with GUI environments, which consume substantial hardware resources as the number of sandboxes scales. Stochastic e…