3 papers
cs.CV2026
On Locality and Length Generalization in Visual Reasoning
Pulkit Madan, Sanjay Haresh, Reza Ebrahimi +3
A striking feature of the human visual system is that it ingests visual information through a series of local foveated glimpses, rather than a single global computation. This makes…
cs.CV2025
Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
Apratim Bhattacharyya, Bicheng Xu, Sanjay Haresh +6
Multi-modal Large Language Models (LLM) have advanced conversational abilities but struggle with providing live, interactive step-by-step guidance, a key capability for future AI a…
cs.CV2025
Can Vision-Language Models Answer Face to Face Questions in the Real-World?
Reza Pourreza, Rishit Dagli, Apratim Bhattacharyya +3
AI models have made significant strides in recent years in their ability to describe and answer questions about real-world images. They have also made progress in the ability to co…