2 citations · 2 across the 15 of their papers we have counts for
7 papers · 1 filter
SMILE: Smooth Motion for Improved Long-Horizon VLA Execution
Jongwoo Park, E-Ro Nguyen, Kanchana Ranasinghe +3
Vision-Language-Action (VLA) models reduce inference cost by executing multiple actions per call, but longer horizons often degrade accuracy because raw chunks contain jitter and o…
LACE: Latent Visual Representation for Cross-Embodiment Learning
Yoo Sung Jang, Kanchana Ranasinghe, Cristina Mata +3
Cross-embodiment learning from human demonstrations is hindered by the visual gap between human and robot embodiments. While self-supervised learning (SSL) backbones encode rich in…
IVRA: Improving Visual-Token Relations for Robot Action Policy with Training-Free Hint-Based Guidance
Jongwoo Park, Kanchana Ranasinghe, Jinhyeok Jang +3
Many Vision-Language-Action (VLA) models flatten image patches into a 1D token sequence, weakening the 2D spatial cues needed for precise manipulation. We introduce IVRA, a lightwe…
Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
Yu Fang, Kanchana Ranasinghe, Le Xue +10
Vision-Language-Action (VLA) models have achieved remarkable progress in robotic manipulation by mapping multimodal observations and instructions directly to actions. However, they…
Pixel Motion Diffusion is What We Need for Robot Control
E-Ro Nguyen, Yichi Zhang, Kanchana Ranasinghe +2
We present DAWN (Diffusion is All We Need for robot control), a unified diffusion-based framework for language-conditioned robotic manipulation that bridges high-level motion inten…
Pixel Motion as Universal Representation for Robot Control
Kanchana Ranasinghe, Xiang Li, E-Ro Nguyen +3
We present LangToMo, a vision-language-action framework structured as a dual-system architecture that uses pixel motion forecasts as intermediate representations. Our high-level Sy…