3 papers
cs.CV2026
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
Martin Q. Ma, Willis Guo, Aditya Agrawal +4
Large vision-language models (VLMs) have advanced multimodal tasks such as video question answering (QA). However, VLMs face the challenge of selecting frames effectively and effic…
cs.CV2026
Act2See: Emergent Active Visual Perception for Video Reasoning
Martin Q. Ma, Yuxiao Qu, Aditya Agrawal +4
Vision-Language Models (VLMs) typically rely on static initial frames for video reasoning, restricting their ability to incorporate essential dynamic information as the reasoning p…
cs.CL2025
CoLoTa: A Dataset for Entity-based Commonsense Reasoning over Long-Tail Knowledge
Armin Toroghi, Willis Guo, Scott Sanner
The rise of Large Language Models (LLMs) has redefined the AI landscape, particularly due to their ability to encode factual and commonsense knowledge, and their outstanding perfor…