3 papers
cs.RO2026
TAP-VLA: Tactile Annotation Prompting for Vision Language Action Models
Mark Van der Merwe, Mohamad Louai Shehab, Jayjun Lee +4
Vision-Language-Action (VLA) models demonstrate impressive reasoning over visual, semantic, and spatial task variations by leveraging large-scale vision and language pre-training.…
cs.RO2025
In-Context Iterative Policy Improvement for Dynamic Manipulation
Mark Van der Merwe, Devesh Jha
Attention-based architectures trained on internet-scale language data have demonstrated state of the art reasoning ability for various language-based tasks, such as logic problems…
cs.RO2025
This&That: Language-Gesture Controlled Video Generation for Robot Planning
Boyang Wang, Nikhil Sridhar, Chao Feng +4
Clear, interpretable instructions are invaluable when attempting any complex task. Good instructions help to clarify the task and even anticipate the steps needed to solve it. In t…