most citedViTacFormer: Learning Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation

1 citations · 1 across the 4 of their papers we have counts for

collaborators

17 papers

cs.RO2026

DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization

Yikai Tang, Haoran Geng, Jindou Jia +5

Imitation learning has emerged as a crucial approach for acquiring visuomotor skills from demonstrations, where designing effective observation encoders is essential for policy gen…

cs.LG2026

Multi-Objective Learning for Diffusion Models: A Statistical Theory under Semi-Supervised Learning

Ziheng Cheng, Yixiao Huang, Hanlin Zhu +5

Diffusion models are increasingly used as powerful conditional generators, yet real deployments often involve multiple target distributions arising from different tasks, e.g., dive…

cs.RO20261 cited

ViTacFormer: Learning Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation

Liang Heng, Haoran Geng, Kaifeng Zhang +2

Dexterous manipulation is a cornerstone capability for robotic systems aiming to interact with the physical world in a human-like manner. Although vision-based methods have advance…

cs.RO2026

Large Video Planner Enables Generalizable Robot Control

Boyuan Chen, Tianyuan Zhang, Haoran Geng +9

General-purpose robots require decision-making models that generalize across diverse tasks and environments. Recent works build robot foundation models by extending multimodal larg…

cs.RO2026

Rodrigues Network for Learning Robot Actions

Jialiang Zhang, Haoran Geng, Yang You +4

Understanding and predicting articulated actions is important in robot learning. However, common architectures such as MLPs and Transformers lack inductive biases that reflect the…

cs.DC2026

OSGym: Scalable OS Infra for Computer Use Agents

Zengyi Qin, Jinyuan Chen, Yunze Man +25

Training computer use agents requires full-featured OS sandboxes with GUI environments, which consume substantial hardware resources as the number of sandboxes scales. Stochastic e…