Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
Martin Q. Ma, Willis Guo, Aditya Agrawal +4
Large vision-language models (VLMs) have advanced multimodal tasks such as video question answering (QA). However, VLMs face the challenge of selecting frames effectively and effic…
cs.CV2026
Act2See: Emergent Active Visual Perception for Video Reasoning
Martin Q. Ma, Yuxiao Qu, Aditya Agrawal +4
Vision-Language Models (VLMs) typically rely on static initial frames for video reasoning, restricting their ability to incorporate essential dynamic information as the reasoning p…