15 papers
Model Confidence Under Answer-Preserving Attacks: An Informativeness-Manipulability Frontier
Reza Khanmohammadi, Ivan Brugere, Simerjot Kaur +3
Deployed vision-language systems often gate their answers on confidence, making confidence robustness relevant to oversight. We study confidence readouts under white-box, image-onl…
Confidence Estimation for Financial Vision-Language Models in Chart and Document Understanding
Reza Khanmohammadi, Simerjot Kaur, Charese H. Smiley +2
LVLMs are increasingly used to read financial charts, tables, and documents, where a single misread figure can move a decision and the most authoritative-looking answer is sometime…
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
Chengzhi Liu, Yuzhe Yang, Sophia Xiao Pu +14
Multimodal large language models are increasingly deployed as long-horizon agents, where memory must do more than recall: it must track an evolving world, revise what has gone stal…
Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking
Reza Khanmohammadi, Erfan Miahi, Simerjot Kaur +4
Large vision-language models suffer from visual ungroundedness: they can produce a fluent, confident, and even correct response driven entirely by language priors, with the image c…
Deep FinResearch Bench: Evaluating AI's Ability to Conduct Professional Financial Investment Research
Mirazul Haque, Antony Papadimitriou, Samuel Mensah +6
We introduce Deep FinResearch Bench, a practical and comprehensive evaluation framework for deep research (DR) agents in financial investment research. The benchmark assesses three…
Distill and Align Decomposition for Enhanced Claim Verification
Jabez Magomere, Elena Kochkina, Samuel Mensah +6
Complex claim verification requires decomposing sentences into verifiable subclaims, yet existing methods struggle to align decomposition quality with verification performance. We…