3 papers
cs.RO2026
CT-VAM: A Cerebello-Thalamic-Inspired Vision-Action Model for Efficient Visuomotor Control
Jiacheng Li, Yize Guo, Jiabin Guo +2
Vision-language-action models have shown strong promise for robot manipulation, yet raw language is primarily needed to specify task intent rather than to be repeatedly processed d…
cs.RO2026
SyncTwin: Fast Digital Twin Construction and Synchronization for Safe Robotic Manipulation
Ruopeng Huang, Boyu Yang, Wenlong Gui +3
Accurate and safe robotic manipulation under dynamic and visually occluded conditions remains a core challenge in real-world deployment. We introduce SyncTwin, a novel digital twin…
cs.RO2025
AutoFocus-IL: VLM-based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations
Litian Gong, Fatemeh Bahrani, Yutai Zhou +3
AutoFocus-IL is a simple yet effective method to improve data efficiency and generalization in visual imitation learning by guiding policies to attend to task-relevant features rat…