4 papers
Pixel Motion Diffusion is What We Need for Robot Control
E-Ro Nguyen, Yichi Zhang, Kanchana Ranasinghe +2
We present DAWN (Diffusion is All We Need for robot control), a unified diffusion-based framework for language-conditioned robotic manipulation that bridges high-level motion inten…
Pixel Motion as Universal Representation for Robot Control
Kanchana Ranasinghe, Xiang Li, E-Ro Nguyen +3
We present LangToMo, a vision-language-action framework structured as a dual-system architecture that uses pixel motion forecasts as intermediate representations. Our high-level Sy…
Understanding Long Videos with Multimodal Language Models
Kanchana Ranasinghe, Xiang Li, Kumara Kahatapitiya +1
Large Language Models (LLMs) have allowed recent LLM-based approaches to achieve excellent performance on long-video understanding benchmarks. We investigate how extensive world kn…
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
Xiang Li, Cristina Mata, Jongwoo Park +8
Vision Language Models (VLMs) have recently been leveraged to generate robotic actions, forming Vision-Language-Action (VLA) models. However, directly adapting a pretrained VLM for…