40 citations · 59 across the 65 of their papers we have counts for
52 papers · 1 filter
Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?
Xin Chen, Dongliang Xu, Cunhao Zhu +5
As vision-language models (VLMs) are increasingly applied to medical AI, existing benchmarks mainly focus on evaluating their diagnostic ability over given medical images and texts…
Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents
Xueqiao Sun, Xiaohan Wang, Ludwig Schmidt +2
Computer-use agents, which leverage multimodal large language models (MLLMs) to operate computers and complete tasks, have attracted significant attention for their utility and ver…
SalArt-VQA: Diagnosing Whether VLMs Understand Salient Artifacts in Generated Images
Xiaoxiao Sun, Ruotian Zhang, Junzhe Huang +2
Vision-language models (VLMs) are increasingly used to detect whether AI-generated images contain visible artifacts, yet their ability to analyze such artifacts remains poorly unde…
MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models
Ryan D'Cunha, Alejandro Lozano, Xiaoxiao Sun +17
Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscop…
DataComp-VLM: Improved Open Datasets for Vision-Language Models
Matteo Farina, Vishaal Udandarao, Thao Nguyen +33
Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curat…
WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild
Junzhe Huang, Xiaoxiao Sun, Yan Yang +6
Using multimodal foundation models to analyze table images is a high-value yet challenging application in consumer and enterprise scenarios. Despite its importance, current evaluat…