activity
20222024
most citedTimeBalance: Temporally-Invariant and Temporally-Distinctive Video Representations for Semi-Supervised Action Recognition

2 citations · 6 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV2024

GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers

Manu S Pillai, Mamshad Nayeem Rizve, Mubarak Shah

Cross-view video geo-localization (CVGL) aims to derive GPS trajectories from street-view videos by aligning them with aerial-view images. Despite their promising performance, curr…

cs.CV20241 cited

X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs

Sirnam Swetha, Jinyu Yang, Tal Neiman +5

Recent advancements in Multimodal Large Language Models (MLLMs) have revolutionized the field of vision-language understanding by integrating visual perception capabilities into La…

cs.CV2024

VidLA: Video-Language Alignment at Scale

Mamshad Nayeem Rizve, Fan Fei, Jayakrishnan Unnikrishnan +5

In this paper, we propose VidLA, an approach for video-language alignment at scale. There are two major limitations of previous video-language alignment approaches. First, they do…

cs.CV2023

CDFSL-V: Cross-Domain Few-Shot Learning for Videos

Sarinda Samarasinghe, Mamshad Nayeem Rizve, Navid Kardan +1

Few-shot video action recognition is an effective approach to recognizing new categories with only a few labeled examples, thereby reducing the challenges associated with collectin…

cs.CV2023

Preserving Modality Structure Improves Multi-Modal Learning

Swetha Sirnam, Mamshad Nayeem Rizve, Nina Shvetsova +2

Self-supervised learning on large-scale multi-modal datasets allows learning semantically meaningful embeddings in a joint multi-modal representation space without relying on human…

cs.CV20232 cited

TimeBalance: Temporally-Invariant and Temporally-Distinctive Video Representations for Semi-Supervised Action Recognition

Ishan Rajendrakumar Dave, Mamshad Nayeem Rizve, Chen Chen +1

Semi-Supervised Learning can be more beneficial for the video domain compared to images because of its higher annotation cost and dimensionality. Besides, any video understanding t…