activity
20242026
collaborators

8 papers

cs.RO2026

IVRA: Improving Visual-Token Relations for Robot Action Policy with Training-Free Hint-Based Guidance

Jongwoo Park, Kanchana Ranasinghe, Jinhyeok Jang +3

Many Vision-Language-Action (VLA) models flatten image patches into a 1D token sequence, weakening the 2D spatial cues needed for precise manipulation. We introduce IVRA, a lightwe…

cs.RO2026

LACE: Latent Visual Representation for Cross-Embodiment Learning

Yoo Sung Jang, Kanchana Ranasinghe, Cristina Mata +3

Cross-embodiment learning from human demonstrations is hindered by the visual gap between human and robot embodiments. While self-supervised learning (SSL) backbones encode rich in…

cs.RO2026

Pixel Motion Diffusion is What We Need for Robot Control

E-Ro Nguyen, Yichi Zhang, Kanchana Ranasinghe +2

We present DAWN (Diffusion is All We Need for robot control), a unified diffusion-based framework for language-conditioned robotic manipulation that bridges high-level motion inten…

cs.CV2026

Phrase-Instance Alignment for Generalized Referring Segmentation

E-Ro Nguyen, Hieu Le, Dimitris Samaras +1

Generalized Referring expressions can describe one object, several related objects, or none at all. Existing generalized referring segmentation (GRES) models treat all cases alike,…

cs.RO2025

Pixel Motion as Universal Representation for Robot Control

Kanchana Ranasinghe, Xiang Li, E-Ro Nguyen +3

We present LangToMo, a vision-language-action framework structured as a dual-system architecture that uses pixel motion forecasts as intermediate representations. Our high-level Sy…

cs.CV2025

Image Translation with Kernel Prediction Networks for Semantic Segmentation

Cristina Mata, Michael S. Ryoo, Henrik Turbell

Semantic segmentation relies on many dense pixel-wise annotations to achieve the best performance, but owing to the difficulty of obtaining accurate annotations for real world data…