8 papers
Imagine How To Change: Explicit Procedure Modeling for Change Captioning
Jiayang Sun, Zixin Guo, Min Cao +2
Change captioning generates descriptions that explicitly describe the differences between two visually similar images. Existing methods operate on static image pairs, thus ignoring…
OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning
Zongyan Han, Jiale Cao, Shuo Chen +3
Open-Vocabulary Segmentation (OVS) has drawn increasing attention for its capacity to generalize segmentation beyond predefined categories. However, existing methods typically pred…
Bilateral Reference for High-Resolution Dichotomous Image Segmentation
Peng Zheng, Dehong Gao, Deng-Ping Fan +4
We introduce a novel bilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS). It comprises two essential components: the localization modu…
MIRA: A Novel Framework for Fusing Modalities in Medical RAG
Jinhong Wang, Tajamul Ashraf, Zongyan Han +2
Multimodal Large Language Models (MLLMs) have significantly advanced AI-assisted medical diagnosis, but they often generate factually inconsistent responses that deviate from estab…
ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark
Sara Ghaboura, Ketan More, Wafa Alghallabi +5
As Large Multimodal Models (LMMs) become more capable, there is growing interest in evaluating their reasoning processes alongside their final outputs. However, most benchmarks rem…
XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models
Omkar Thawakar, Abdelrahman Shaker, Sahal Shaji Mullappilly +5
The latest breakthroughs in large vision-language models, such as Bard and GPT-4, have showcased extraordinary abilities in performing a wide range of tasks. Such models are traine…