8 citations · 8 across the 5 of their papers we have counts for
10 papers · 1 filter
Imagine How To Change: Explicit Procedure Modeling for Change Captioning
Jiayang Sun, Zixin Guo, Min Cao +2
Change captioning generates descriptions that explicitly describe the differences between two visually similar images. Existing methods operate on static image pairs, thus ignoring…
MIRA: A Novel Framework for Fusing Modalities in Medical RAG
Jinhong Wang, Tajamul Ashraf, Zongyan Han +2
Multimodal Large Language Models (MLLMs) have significantly advanced AI-assisted medical diagnosis, but they often generate factually inconsistent responses that deviate from estab…
ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark
Sara Ghaboura, Ketan More, Wafa Alghallabi +5
As Large Multimodal Models (LMMs) become more capable, there is growing interest in evaluating their reasoning processes alongside their final outputs. However, most benchmarks rem…
OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning
Zongyan Han, Jiale Cao, Shuo Chen +3
Open-Vocabulary Segmentation (OVS) has drawn increasing attention for its capacity to generalize segmentation beyond predefined categories. However, existing methods typically pred…
All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages
Ashmal Vayani, Dinura Dissanayake, Hasindri Watawana +66
Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cul…
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark
Sara Ghaboura, Ahmed Heakl, Omkar Thawakar +7
Recent years have witnessed a significant interest in developing large multimodal models (LMMs) capable of performing various visual reasoning and understanding tasks. This has led…