collaborators

7 papers

cs.CR2026

Model Confidence Under Answer-Preserving Attacks: An Informativeness-Manipulability Frontier

Reza Khanmohammadi, Ivan Brugere, Simerjot Kaur +3

Deployed vision-language systems often gate their answers on confidence, making confidence robustness relevant to oversight. We study confidence readouts under white-box, image-onl…

cs.CL2026

Confidence Estimation for Financial Vision-Language Models in Chart and Document Understanding

Reza Khanmohammadi, Simerjot Kaur, Charese H. Smiley +2

LVLMs are increasingly used to read financial charts, tables, and documents, where a single misread figure can move a decision and the most authoritative-looking answer is sometime…

cs.CL2026

Calibrated Triage, Not Autonomy: Confidence Estimation for Medical Vision-Language Models

Reza Khanmohammadi, Kundan Thind, Mohammad M. Ghassemi

A vision-language model can answer a question about a chest radiograph or a pathology slide fluently and confidently while barely using the image, relying instead on language prior…

cs.CL2026

Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking

Reza Khanmohammadi, Erfan Miahi, Simerjot Kaur +4

Large vision-language models suffer from visual ungroundedness: they can produce a fluent, confident, and even correct response driven entirely by language priors, with the image c…

cs.CL2026

How Reliable are Confidence Estimators for Large Reasoning Models? A Systematic Benchmark on High-Stakes Domains

Reza Khanmohammadi, Erfan Miahi, Simerjot Kaur +4

The miscalibration of Large Reasoning Models (LRMs) undermines their reliability in high-stakes domains, necessitating methods to accurately estimate the confidence of their long-f…

cs.AI2025

Automated stereotactic radiosurgery planning using a human-in-the-loop reasoning large language model agent

Humza Nusrat, Luke Francisco, Bing Luo +10

Stereotactic radiosurgery (SRS) demands precise dose shaping around critical structures, yet black-box AI systems have limited clinical adoption due to opacity concerns. We tested…