6 papers
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
Dominick Reilly, Manish Kumar Govind, Le Xue +1
Large Vision Language Models (VLMs) excel at general visual reasoning but experience significant performance degradation when deployed in novel domains that exhibit substantial dis…
From My View to Yours: Learning Egocentric Cues from Exocentric Video using Privileged Egocentric Supervision
Dominick Reilly, Manish Kumar Govind, Le Xue +1
Vision Language Models (VLMs) have achieved strong performance across a wide range of video understanding tasks. However, their viewpoint-invariant training limits their ability to…
World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays
Manish Kumar Govind, Dominick Reilly, Smit Patel +2
Going beyond predicting robot actions, World Action Models (WAMs) can also generate future visual observations. We build on this generative capability to propose Recurrent Generati…
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
Manish Kumar Govind, Dominick Reilly, Pu Wang +1
Latent action representations learned from unlabeled videos have recently emerged as a promising paradigm for pretraining vision-language-action (VLA) models without explicit robot…
Fibottention: Inceptive Visual Representation Learning with Diverse Attention Across Heads
Ali K. Rahimian, Manish K. Govind, Subhajit Maity +4
Vision Transformers and their variants have achieved remarkable success in diverse visual perception tasks. Despite their effectiveness, they suffer from two significant limitation…
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
Dominick Reilly, Rajatsubhra Chakraborty, Arkaprava Sinha +5
Current Large Language Vision Models (LLVMs) trained on web videos perform well in general video understanding but struggle with fine-grained details, complex human-object interact…