4 papers
Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence
Sankalp Nagaonkar, Rohit Garg, Ankit Raj +2
Most video-retrieval systems assume a bounded corpus and return ranked files or timestamps. Agents operating over cameras, screens, streams, and archives face a different systems p…
Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding
Shivam Sharma, Sankalp Nagaonkar, Ashish Choithani +1
We benchmark how internal reasoning traces, which we call thought streams, affect video scene understanding in vision-language models. Using four configurations of Google's Gemini…
Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments
Sankalp Nagaonkar, Augustya Sharma, Ashish Choithani +1
This paper introduces an open-source benchmark for evaluating Vision-Language Models (VLMs) on Optical Character Recognition (OCR) tasks in dynamic video environments. We present a…
BadScan: An Architectural Backdoor Attack on Visual State Space Models
Om Suhas Deshmukh, Sankalp Nagaonkar, Achyut Mani Tripathi +1
The newly introduced Visual State Space Model (VMamba), which employs \textit{State Space Mechanisms} (SSM) to interpret images as sequences of patches, has shown exceptional perfo…