activity
20222026
most citedPiTL: Cross-modal Retrieval with Weakly-supervised Vision-language Pre-training via Prompting

8 citations · 8 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2026

Imagine How To Change: Explicit Procedure Modeling for Change Captioning

Jiayang Sun, Zixin Guo, Min Cao +2

Change captioning generates descriptions that explicitly describe the differences between two visually similar images. Existing methods operate on static image pairs, thus ignoring…

cs.CV2025

MIRA: A Novel Framework for Fusing Modalities in Medical RAG

Jinhong Wang, Tajamul Ashraf, Zongyan Han +2

Multimodal Large Language Models (MLLMs) have significantly advanced AI-assisted medical diagnosis, but they often generate factually inconsistent responses that deviate from estab…

cs.CV2025

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark

Sara Ghaboura, Ketan More, Wafa Alghallabi +5

As Large Multimodal Models (LMMs) become more capable, there is growing interest in evaluating their reasoning processes alongside their final outputs. However, most benchmarks rem…

cs.CV2025

OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning

Zongyan Han, Jiale Cao, Shuo Chen +3

Open-Vocabulary Segmentation (OVS) has drawn increasing attention for its capacity to generalize segmentation beyond predefined categories. However, existing methods typically pred…

cs.CV20241 cited

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

Ashmal Vayani, Dinura Dissanayake, Hasindri Watawana +66

Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cul…

cs.CV2024

CAMEL-Bench: A Comprehensive Arabic LMM Benchmark

Sara Ghaboura, Ahmed Heakl, Omkar Thawakar +7

Recent years have witnessed a significant interest in developing large multimodal models (LMMs) capable of performing various visual reasoning and understanding tasks. This has led…