9 papers
Personal Visual Context Learning in Large Multimodal Models
Zihui Xue, Ami Baid, Sangho Kim +2
As wearable devices like smart glasses integrate Large Multimodal Models (LMMs) into the continuous first-person visual streams of individual users, the evolution of these models i…
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
Ami Baid, Zihui Xue, Kristen Grauman
While Audio-Visual Language Models (AVLMs) have achieved remarkable progress over recent years, their reliability is bottlenecked by cross-modal hallucination. A particularly perva…
Seeing without Pixels: Perception from Camera Trajectories
Zihui Xue, Kristen Grauman, Dima Damen +2
Can one perceive a video's content without seeing its pixels, just from the camera trajectory-the path it carves through space? This paper is the first to systematically investigat…
SPOC: Spatially-Progressing Object State Change Segmentation in Video
Priyanka Mandikal, Tushar Nagarajan, Alex Stoken +2
Object state changes in video reveal critical cues about human and agent activity. However, existing methods are limited to temporal localization of when the object is in its initi…
Seeing the Arrow of Time in Large Multimodal Models
Zihui Xue, Mi Luo, Kristen Grauman
The Arrow of Time (AoT)-time's irreversible flow shaping physical events-is fundamental to video comprehension, yet remains a significant challenge for modern large multimodal mode…
When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
Mi Luo, Zihui Xue, Alex Dimakis +1
Video reasoning, the task of enabling machines to infer from dynamic visual content through multi-step logic, is crucial for advanced AI. While the Chain-of-Thought (CoT) mechanism…