10 papers
Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing
Dominick Reilly, Qiyu Wu, Hiromi Wakaki +2
Multimodal Large Language Models (MLLMs) are typically designed under the assumption that all modalities available during training will also be accessible at inference. However, ma…
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
Dominick Reilly, Manish Kumar Govind, Le Xue +1
Large Vision Language Models (VLMs) excel at general visual reasoning but experience significant performance degradation when deployed in novel domains that exhibit substantial dis…
From My View to Yours: Learning Egocentric Cues from Exocentric Video using Privileged Egocentric Supervision
Dominick Reilly, Manish Kumar Govind, Le Xue +1
Vision Language Models (VLMs) have achieved strong performance across a wide range of video understanding tasks. However, their viewpoint-invariant training limits their ability to…
World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays
Manish Kumar Govind, Dominick Reilly, Smit Patel +2
Going beyond predicting robot actions, World Action Models (WAMs) can also generate future visual observations. We build on this generative capability to propose Recurrent Generati…
TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living
Arkaprava Sinha, Dominick Reilly, Siddharth Krishnan +2
Long Video Question Answering (LVQA) requires identifying sparse, query-relevant evidence within hours-long untrimmed videos. Existing approaches either process videos densely with…
UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning
Wenhao Chi, Arkaprava Sinha, Dominick Reilly +2
Egocentric video understanding is inherently limited by the narrow perspective of wearable cameras: a single viewpoint, a single modality, a single model cannot capture the full ri…