Publications (5)
CompCap: Improving Multimodal Large Language Models with Composite Captions
Xiaohui Chen, Satya Narayan Shukla, Mahmoud Azab +8
How well can Multimodal Large Language Models (MLLMs) understand composite images? Composite images (CIs) are synthetic visuals created by merging multiple visual elements, such as…
A Comparative Analysis of Content-based Geolocation in Blogs and Tweets
Konstantinos Pappas, Mahmoud Azab, Rada Mihalcea
The geolocation of online information is an essential component in any geospatial application. While most of the previous work on geolocation has focused on Twitter, in this paper…
Fighting FIRe with FIRE: Assessing the Validity of Text-to-Video Retrieval Benchmarks
Pedro Rodriguez, Mahmoud Azab, Becka Silvert +4
Searching troves of videos with textual descriptions is a core multimodal retrieval task. Owing to the lack of a purpose-built dataset for text-to-video retrieval, video captioning…
Speaker Naming in Movies
Mahmoud Azab, Mingzhe Wang, Max Smith +3
We propose a new model for speaker naming in movies that leverages visual, textual, and acoustic modalities in an unified optimization framework. To evaluate the performance of our…
Normalized Contrastive Learning for Text-Video Retrieval
Yookoon Park, Mahmoud Azab, Bo Xiong +4
Cross-modal contrastive learning has led the recent advances in multimodal retrieval with its simplicity and effectiveness. In this work, however, we reveal that cross-modal contra…