2 papers
cs.CV2026
Incentivizing Vision Language Models to Search for Long Video Question Answering
Harsh Goel, S P Sharan, Sahil Shah +4
We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process. VSeek…
cs.CV2026
UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding
Joungbin An, Agrim Jain, Kristen Grauman
Video temporal grounding (VTG) is typically tackled with dataset-specific models that transfer poorly across domains and query styles. Recent efforts to overcome this limitation ha…