activity
20232026
most citedTowards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation

40 citations · 59 across the 65 of their papers we have counts for

collaborators
Showing cs.CVShow all

52 papers · 1 filter

cs.CV2026

Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?

Xin Chen, Dongliang Xu, Cunhao Zhu +5

As vision-language models (VLMs) are increasingly applied to medical AI, existing benchmarks mainly focus on evaluating their diagnostic ability over given medical images and texts…

cs.CV2026

Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents

Xueqiao Sun, Xiaohan Wang, Ludwig Schmidt +2

Computer-use agents, which leverage multimodal large language models (MLLMs) to operate computers and complete tasks, have attracted significant attention for their utility and ver…

cs.CV2026

SalArt-VQA: Diagnosing Whether VLMs Understand Salient Artifacts in Generated Images

Xiaoxiao Sun, Ruotian Zhang, Junzhe Huang +2

Vision-language models (VLMs) are increasingly used to detect whether AI-generated images contain visible artifacts, yet their ability to analyze such artifacts remains poorly unde…

cs.CV2026

MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models

Ryan D'Cunha, Alejandro Lozano, Xiaoxiao Sun +17

Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscop…

cs.CV2026

DataComp-VLM: Improved Open Datasets for Vision-Language Models

Matteo Farina, Vishaal Udandarao, Thao Nguyen +33

Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curat…

cs.CV2026

WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild

Junzhe Huang, Xiaoxiao Sun, Yan Yang +6

Using multimodal foundation models to analyze table images is a high-value yet challenging application in consumer and enterprise scenarios. Despite its importance, current evaluat…