1 citations · 1 across the 4 of their papers we have counts for
5 papers
Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos
Bonan Ding, Umair Nawaz, Ufaq Khan +5
Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, depth map, or dense temporal evi…
CoVR-R:Reason-Aware Composed Video Retrieval
Omkar Thawakar, Dmitry Demidov, Vaishnav Potlapalli +5
Composed Video Retrieval (CoVR) aims to find a target video given a reference video and a textual modification. Prior work assumes the modification text fully specifies the visual…
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
Shiming Chen, Bowen Duan, Salman Khan +1
Large-scale vision-language models (VLMs), such as CLIP, have achieved remarkable success in zero-shot learning (ZSL) by leveraging large-scale visual-text pair datasets. However,…
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
Abdelrahman Shaker, Muhammad Maaz, Chenhui Gou +3
Video understanding models often struggle with high computational requirements, extensive parameter counts, and slow inference speed, making them inefficient for practical use. To…
ZeroDiff: Solidified Visual-Semantic Correlation in Zero-Shot Learning
Zihan Ye, Shreyank N. Gowda, Xiaowei Huang +4
Zero-shot Learning (ZSL) aims to enable classifiers to identify unseen classes. This is typically achieved by generating visual features for unseen classes based on learned visual-…