1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 1 cited
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
Sirnam Swetha, Jinyu Yang, Tal Neiman +5
Recent advancements in Multimodal Large Language Models (MLLMs) have revolutionized the field of vision-language understanding by integrating visual perception capabilities into La…
cs.CV2024
VidLA: Video-Language Alignment at Scale
Mamshad Nayeem Rizve, Fan Fei, Jayakrishnan Unnikrishnan +5
In this paper, we propose VidLA, an approach for video-language alignment at scale. There are two major limitations of previous video-language alignment approaches. First, they do…
cs.CV2023
UnsMOT: Unified Framework for Unsupervised Multi-Object Tracking with Geometric Topology Guidance
Son Tran, Cong Tran, Anh Tran +1
Object detection has long been a topic of high interest in computer vision literature. Motivated by the fact that annotating data for the multi-object tracking (MOT) problem is imm…