27 papers
Incentivizing Vision Language Models to Search for Long Video Question Answering
Harsh Goel, S P Sharan, Sahil Shah +4
We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process. VSeek…
Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
Daniel Adebi, Sagnik Majumder, Kristen Grauman
Understanding camera motion is a fundamental problem in embodied perception and 3D scene understanding. While visual methods have advanced rapidly, they often struggle under visual…
EgoExo-WM: Unlocking Exo Video for Ego World Models
Danny Tran, Roberto MartÃn-MartÃn, Kristen Grauman
Egocentric world models present a promising direction for enabling agents to predict and plan, but their performance is constrained by the limited availability of egocentric traini…
ViewBridge: Curriculum Knowledge Distillation for Activity View-Invariance Under Extreme Viewpoint Changes
Arjun Somayazulu, Efi Mavroudi, Changan Chen +2
Traditional methods for view-invariant learning rely on controlled multi-view training data with minimal scene clutter. However, they struggle with in-the-wild videos that exhibit…
Personal Visual Context Learning in Large Multimodal Models
Zihui Xue, Ami Baid, Sangho Kim +2
As wearable devices like smart glasses integrate Large Multimodal Models (LMMs) into the continuous first-person visual streams of individual users, the evolution of these models i…
Materialistic RIR: Material Conditioned Realistic RIR Generation
Mahnoor Fatima Saad, Sagnik Majumder, Kristen Grauman +1
Rings like gold, thuds like wood! The sound we hear in a scene is shaped not only by the spatial layout of the environment but also by the materials of the objects and surfaces wit…