5 papers
The Embodiment Gap in Robot Foundation Models
Yukiyasu Domae, Keisuke Shirai, Hanbit Oh +7
Robot foundation models (RFMs), including vision-language-action (VLA) policies, are often discussed through a scaling view: more data, larger models, and broader benchmarks should…
GuidedAttention: Interpretable and Correctable Visual Attention for OOD-Robust Robot Manipulation via Imitation Learning
Masaki Murooka, Ryoichi Nakajo, Keisuke Shirai +4
End-to-end visuomotor policies provide little opportunity for humans to understand or correct the policy's visual attention. We propose GuidedAttention, a visuomotor imitation lear…
RoboManipBaselines: A Unified Framework for Imitation Learning in Robotic Manipulation across Real and Simulation Environments
Masaki Murooka, Tomohiro Motoda, Ryoichi Nakajo +5
We present RoboManipBaselines, an open-source software framework for imitation learning research in robotic manipulation. The framework supports the entire imitation learning pipel…
Grounded Vision-Language Interpreter for Long-Horizon Bimanual Task and Motion Planning
Jeremy Siburian, Keisuke Shirai, Cristian C. Beltran-Hernandez +3
While recent advances in vision-language models have accelerated language-guided robot planning, their black-box nature lacks the safety guarantees and interpretability crucial for…
KeyMPs: One-Shot Vision-Language Guided Motion Generation by Sequencing DMPs for Occlusion-Rich Tasks
Edgar Anarossi, Yuhwan Kwon, Hirotaka Tahara +6
Dynamic Movement Primitives (DMPs) provide a flexible framework wherein smooth robotic motions are encoded into modular parameters. However, they face challenges in integrating mul…