2 papers
cs.CV2026
Evaluating the Diagnostic Robustness of Vision-Language Models Under Visual and Textual Perturbations
Ali Khoramfar, Mohammad Javad Dousti, Alireza Mohamadian +1
Standard accuracy metrics for VLMs often mask significant reliability failures in sensitive domains. In this work, we utilize a histopathology-validated brain MRI dataset to system…
cs.CL2026
DeepQuestion: Systematic Generation of Real-World Challenges for Evaluating LLMs Performance
Ali Khoramfar, Ali Ramezani, Mohammad Mahdi Mohajeri +3
While Large Language Models (LLMs) achieve near-human performance on standard benchmarks, their capabilities often fail to generalize to complex, real-world problems. To bridge thi…