3 papers
cs.RO2026
CLAP: Direct VLM-to-VLA Adaptation via Language-Action Grounding
Yuri Ishitoya, Jeremy Siburian, Masashi Hamaya +3
Vision-language-action models (VLAs) inherit semantic capabilities from pretrained VLMs, yet large-scale post-training on robot data and architectural modifications can reshape the…
cs.RO2026
Tactile Memory with Soft Robot: Robust Object Insertion via Masked Encoding and Soft Wrist
Tatsuya Kamijo, Mai Nishimura, Cristian C. Beltran-Hernandez +2
Tactile memory, the ability to store and retrieve touch-based experience, is critical for contact-rich tasks such as key insertion under uncertainty. To replicate this capability,…
cs.RO2025
Grounded Vision-Language Interpreter for Long-Horizon Bimanual Task and Motion Planning
Jeremy Siburian, Keisuke Shirai, Cristian C. Beltran-Hernandez +3
While recent advances in vision-language models have accelerated language-guided robot planning, their black-box nature lacks the safety guarantees and interpretability crucial for…