Showing cs.ROShow all
3 papers · 1 filter
cs.RO2026
Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks
Sanjay Haresh, Daniel Dijkman, Apratim Bhattacharyya +1
Many dexterous manipulation tasks are non-markovian in nature, yet little attention has been paid to this fact in the recent upsurge of the vision-language-action (VLA) paradigm. A…
cs.RO2025
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
Rokas Bendikas, Daniel Dijkman, Markus Peschl +2
Vision-Language-Action (VLA) models offer a pivotal approach to learning robotic manipulation at scale by repurposing large pre-trained Vision-Language-Models (VLM) to output robot…
cs.RO2024
ClevrSkills: Compositional Language and Visual Reasoning in Robotics
Sanjay Haresh, Daniel Dijkman, Apratim Bhattacharyya +1
Robotics tasks are highly compositional by nature. For example, to perform a high-level task like cleaning the table a robot must employ low-level capabilities of moving the effect…