activity
20162026
most citedFoundational Models Defining a New Era in Vision: A Survey and Outlook

68 citations · 336 across the 77 of their papers we have counts for

collaborators
Showing 2024 · cs.CVShow all

14 papers · 2 filters

cs.CV2024

BiMediX2: Bio-Medical EXpert LMM for Diverse Medical Modalities

Sahal Shaji Mullappilly, Mohammed Irfan Kurpath, Sara Pieri +8

We introduce BiMediX2, a bilingual (Arabic-English) Bio-Medical EXpert Large Multimodal Model that supports text-based and image-based medical interactions. It enables multi-turn c…

cs.CV2024★ 1 cited

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

Ashmal Vayani, Dinura Dissanayake, Hasindri Watawana +66

Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cul…

cs.CV2024

CAMEL-Bench: A Comprehensive Arabic LMM Benchmark

Sara Ghaboura, Ahmed Heakl, Omkar Thawakar +7

Recent years have witnessed a significant interest in developing large multimodal models (LMMs) capable of performing various visual reasoning and understanding tasks. This has led…

cs.CV2024

CONDA: Condensed Deep Association Learning for Co-Salient Object Detection

Long Li, Nian Liu, Dingwen Zhang +6

Inter-image association modeling is crucial for co-salient object detection. Despite satisfactory performance, previous methods still have limitations on sufficient inter-image ass…

cs.CV2024★ 2 cited

AgriCLIP: Adapting CLIP for Agriculture and Livestock via Domain-Specialized Cross-Model Alignment

Umair Nawaz, Muhammad Awais, Hanan Gani +4

Capitalizing on vast amount of image-text data, large-scale vision-language pre-training has demonstrated remarkable zero-shot capabilities and has been utilized in several applica…

cs.CV2024

AgroGPT: Efficient Agricultural Vision-Language Model with Expert Tuning

Muhammad Awais, Ali Husain Salem Abdulla Alharthi, Amandeep Kumar +2

Significant progress has been made in advancing large multimodal conversational models (LMMs), capitalizing on vast repositories of image-text data available online. Despite this p…