6 papers
On Locality and Length Generalization in Visual Reasoning
Pulkit Madan, Sanjay Haresh, Reza Ebrahimi +3
A striking feature of the human visual system is that it ingests visual information through a series of local foveated glimpses, rather than a single global computation. This makes…
Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur?
Apratim Bhattacharyya, Shweta Mahajan, Sanjay Haresh +5
Learning everyday skills, like cooking a dish, relies increasingly on instructional media such as online videos. This opens the door to the use of video (and multimodal) large lang…
Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks
Sanjay Haresh, Daniel Dijkman, Apratim Bhattacharyya +1
Many dexterous manipulation tasks are non-markovian in nature, yet little attention has been paid to this fact in the recent upsurge of the vision-language-action (VLA) paradigm. A…
Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
Apratim Bhattacharyya, Bicheng Xu, Sanjay Haresh +6
Multi-modal Large Language Models (LLM) have advanced conversational abilities but struggle with providing live, interactive step-by-step guidance, a key capability for future AI a…
Delayed Attention Training Improves Length Generalization in Transformer--RNN Hybrids
Buu Phan, Reza Ebrahimi, Sanjay Haresh +1
We study length generalization in sequence models on a composite problem involving both state tracking and associative recall. Prior work finds that recurrent networks handle state…
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
Rokas Bendikas, Daniel Dijkman, Markus Peschl +2
Vision-Language-Action (VLA) models offer a pivotal approach to learning robotic manipulation at scale by repurposing large pre-trained Vision-Language-Models (VLM) to output robot…