3 papers
cs.RO2026
IVRA: Improving Visual-Token Relations for Robot Action Policy with Training-Free Hint-Based Guidance
Jongwoo Park, Kanchana Ranasinghe, Jinhyeok Jang +3
Many Vision-Language-Action (VLA) models flatten image patches into a 1D token sequence, weakening the 2D spatial cues needed for precise manipulation. We introduce IVRA, a lightwe…
cs.RO2026
LACE: Latent Visual Representation for Cross-Embodiment Learning
Yoo Sung Jang, Kanchana Ranasinghe, Cristina Mata +3
Cross-embodiment learning from human demonstrations is hindered by the visual gap between human and robot embodiments. While self-supervised learning (SSL) backbones encode rich in…
cs.RO2025
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
Xiang Li, Cristina Mata, Jongwoo Park +8
Vision Language Models (VLMs) have recently been leveraged to generate robotic actions, forming Vision-Language-Action (VLA) models. However, directly adapting a pretrained VLM for…