4 papers · 1 filter
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
Dominick Reilly, Manish Kumar Govind, Le Xue +1
Large Vision Language Models (VLMs) excel at general visual reasoning but experience significant performance degradation when deployed in novel domains that exhibit substantial dis…
From My View to Yours: Learning Egocentric Cues from Exocentric Video using Privileged Egocentric Supervision
Dominick Reilly, Manish Kumar Govind, Le Xue +1
Vision Language Models (VLMs) have achieved strong performance across a wide range of video understanding tasks. However, their viewpoint-invariant training limits their ability to…
Fibottention: Inceptive Visual Representation Learning with Diverse Attention Across Heads
Ali K. Rahimian, Manish K. Govind, Subhajit Maity +4
Vision Transformers and their variants have achieved remarkable success in diverse visual perception tasks. Despite their effectiveness, they suffer from two significant limitation…
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
Dominick Reilly, Rajatsubhra Chakraborty, Arkaprava Sinha +5
Current Large Language Vision Models (LLVMs) trained on web videos perform well in general video understanding but struggle with fine-grained details, complex human-object interact…