2 papers
cs.CV2026
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning
Awais Rauf, Ahmed Hasssan, Greg Slabaugh
Understanding long videos requires fine-grained perception and multi-step, higher-order reasoning over complex, long-range spatio-temporal dynamics. Vision-language models (VLMs) e…
cs.CV2026
BlindSight: Harnessing Sparsity for Efficient Vision-Language Models
Tharun Adithya Srikrishnan, Deval Shah, Timothy Hein +3
Large vision-language models (VLMs) enable joint processing of text and images. However, incorporating vision data significantly increases the prompt length, resulting in a longer…