2 papers
cs.CV2026
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
Han Li, Si Liu, Zehao Huang +6
Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturall…
cs.CV2025
MedGEN-Bench: A Contextually Entangled Benchmark for Open-ended Multimodal Medical Generation
Junjie Yang, Yuhao Yan, Gang Wu +12
Medical vision-language models (VLMs) are increasingly expected to support clinical workflows through diagnostic text and relevant medical images. However, current medical visual b…