2 papers
cs.CV2026
Multimodal Fusion for Sim2real Transfer in Visual Reinforcement Learning
Zichun Xu, Jingdong Zhao, Chenyu Guo +6
Depth information is robust to scene appearance variations and inherently carries 3D spatial details. Thus, a visual backbone based on the vision transformer is proposed to fuse RG…
cs.RO2026
A Visual Reinforcement Learning-Based Separate Primitive Policy for Peg-in-Hole Tasks
Zichun Xu, Zhaomin Wang, Yuntao Li +4
For peg-in-hole tasks, humans rely on binocular visual perception to locate the peg above the hole surface and then proceed with insertion. This paper draws insights from this beha…