activity
20242026
most citedLMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology

1 citations · 1 across the 3 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV20261 cited

LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology

Zhenyue Qin, Yang Liu, Yu Yin +13

Vision-threatening eye diseases pose a major global health burden, with timely diagnosis limited by workforce shortages and restricted access to specialized care. While multimodal…

cs.CV2025

LEMoN: Label Error Detection using Multimodal Neighbors

Haoran Zhang, Aparna Balagopalan, Nassim Oufattole +4

Large repositories of image-caption pairs are essential for the development of vision-language models. However, these datasets are often extracted from noisy data scraped from the…

cs.CV2024

Towards Vision Mixture of Experts for Wildlife Monitoring on the Edge

Emmanuel Azuh Mensah, Anderson Lee, Haoran Zhang +2

The explosion of IoT sensors in industrial, consumer and remote sensing use cases has come with unprecedented demand for computing infrastructure to transmit and to analyze petabyt…

cs.CV2024

Curriculum Prompting Foundation Models for Medical Image Segmentation

Xiuqi Zheng, Yuhang Zhang, Haoran Zhang +4

Adapting large pre-trained foundation models, e.g., SAM, for medical image segmentation remains a significant challenge. A crucial step involves the formulation of a series of spec…

cs.CV2024

Multimodal Information Interaction for Medical Image Segmentation

Xinxin Fan, Lin Liu, Haoran Zhang

The use of multimodal data in assisted diagnosis and segmentation has emerged as a prominent area of interest in current research. However, one of the primary challenges is how to…