4 papers · 1 filter
SPIKE-RL: Video-LLMs meet Bayesian Surprise
Sahithya Ravi, Aditya Chinchure, Raymond T. Ng +2
Real-world videos often show routine activities punctuated by memorable, surprising events. However, most Video-LLMs process videos by sampling frames uniformly, likely missing cri…
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
Sahithya Ravi, Gabriel Sarch, Vibhav Vineet +2
An embodied AI assistant operating on egocentric video must integrate spatial cues across time - for instance, determining where an object A, glimpsed a few moments ago lies relati…
Grounding Task Assistance with Multimodal Cues from a Single Demonstration
Gabriel Sarch, Balasaravanan Thoravi Kumaravel, Sahithya Ravi +2
A person's demonstration often serves as a key reference for others learning the same task. However, RGB video, the dominant medium for representing these demonstrations, often fai…
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events
Aditya Chinchure, Sahithya Ravi, Raymond Ng +3
The commonsense reasoning capabilities of vision-language models (VLMs), especially in abductive reasoning and defeasible reasoning, remain poorly understood. Most benchmarks focus…