2 papers
cs.CV2026
How Do Medical MLLMs Fail? A Study on Visual Grounding in Medical Images
Guimeng Liu, Tianze Yu, Somayeh Ebrahimkhani +3
Generalist multimodal large language models (MLLMs) have achieved impressive performance across a wide range of vision-language tasks. However, their performance on medical tasks,…
cs.LG2025
Vision Transformer Neural Architecture Search for Out-of-Distribution Generalization: Benchmark and Insights
Sy-Tuyen Ho, Tuan Van Vo, Somayeh Ebrahimkhani +1
While ViTs have achieved across machine learning tasks, deploying them in real-world scenarios faces a critical challenge: generalizing under OoD shifts. A crucial research gap exi…