6 papers
Attention from Action, for Action: Emergent Visual Bottlenecks for Policy Learning
Zheyu Zhuang, Ruiyu Wang, Nick Heppert +4
Visual bottlenecks that focus policy inputs on regions of interest (ROIs) can improve data-efficient visuomotor learning by separating where to look from how to act. Many ROI inter…
Enhancing Visual Domain Robustness in Behaviour Cloning via Saliency-Guided Augmentation
Zheyu Zhuang, Ruiyu Wang, Nils Ingelhag +2
In vision-based behavior cloning (BC), conventional image augmentations such as Random Crop and Color Jitter often fall short under substantial visual domain shifts, including chan…
MirrorDuo: Reflection-Consistent Visuomotor Learning from Mirrored Demonstration Pairs
Zheyu Zhuang, Ruiyu Wang, Giovanni Luca Marchetti +2
Image-based behaviour cloning leverages demonstrations captured from ubiquitous RGB cameras. However, it remains constrained by the cost of collecting diverse demos, especially for…
PALM: Enhanced Generalizability for Local Visuomotor Policies via Perception Alignment
Ruiyu Wang, Zheyu Zhuang, Danica Kragic +1
Generalizing beyond the training domain in image-based behavior cloning remains challenging. Existing methods address individual axes of generalization, workspace shifts, viewpoint…
Raising Body Ownership in End-to-End Visuomotor Policy Learning via Robot-Centric Pooling
Zheyu Zhuang, Ville Kyrki, Danica Kragic
We present Robot-centric Pooling (RcP), a novel pooling method designed to enhance end-to-end visuomotor policies by enabling differentiation between the robots and similar entitie…
Feature Extractor or Decision Maker: Rethinking the Role of Visual Encoders in Visuomotor Policies
Ruiyu Wang, Zheyu Zhuang, Shutong Jin +3
An end-to-end (E2E) visuomotor policy is typically treated as a unified whole, but recent approaches using out-of-domain (OOD) data to pretrain the visual encoder have cleanly sepa…