3 papers
cs.CV2024
CompCap: Improving Multimodal Large Language Models with Composite Captions
Xiaohui Chen, Satya Narayan Shukla, Mahmoud Azab +8
How well can Multimodal Large Language Models (MLLMs) understand composite images? Composite images (CIs) are synthetic visuals created by merging multiple visual elements, such as…
cs.CL2018
A Comparative Analysis of Content-based Geolocation in Blogs and Tweets
Konstantinos Pappas, Mahmoud Azab, Rada Mihalcea
The geolocation of online information is an essential component in any geospatial application. While most of the previous work on geolocation has focused on Twitter, in this paper…
cs.CL2018
Speaker Naming in Movies
Mahmoud Azab, Mingzhe Wang, Max Smith +3
We propose a new model for speaker naming in movies that leverages visual, textual, and acoustic modalities in an unified optimization framework. To evaluate the performance of our…