5 papers
Beyond Semantics: Modeling Factual and Affective Perceptual Experiences from Vision-Language Data
Youssef Mohamed, Kenneth Ward Church, Mohamed Elhoseiny
We present P-Topics (Perception Topics) modeling, a novel problem for understanding how images are perceived affectively and across cultures. The goal is to (1) discover and model…
M-MiniGPT4: Multilingual VLLM Alignment via Translated Data
Seung Hun Han, Youssef Mohamed, Mohamed Elhoseiny
This paper presents a Multilingual Vision Large Language Model, named M-MiniGPT4. Our model exhibits strong vision-language understanding (VLU) capabilities across 11 languages. We…
XProvence: Zero-Cost Multilingual Context Pruning for Retrieval-Augmented Generation
Youssef Mohamed, Mohamed Elhoseiny, Thibault Formal +1
This paper introduces XProvence, a multilingual zero-cost context pruning model for retrieval-augmented generation (RAG), trained on 16 languages and supporting 100+ languages thro…
Neural Catalog: Scaling Species Recognition with Catalog of Life-Augmented Generation
Faizan Farooq Khan, Jun Chen, Youssef Mohamed +2
Open-vocabulary species recognition is a major challenge in computer vision, particularly in ornithology, where new taxa are continually discovered. While benchmarks like CUB-200-2…
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
Youssef Mohamed, Runjia Li, Ibrahim Said Ahmad +4
Research in vision and language has made considerable progress thanks to benchmarks such as COCO. COCO captions focused on unambiguous facts in English; ArtEmis introduced subjecti…