3 papers
cs.CV2026
CheXthought: A global multimodal dataset of clinical chain-of-thought reasoning and visual attention for chest X-ray interpretation
Sonali Sharma, Jin Long, George Shih +7
Chest X-ray interpretation is one of the most frequently performed diagnostic tasks in medicine and a primary target for AI development, yet current vision-language models are prim…
cs.CL2026
Advances in LLM Reasoning Enable Flexibility in Clinical Problem-Solving
Kie Shidara, Preethi Prem, Jonathan Kim +4
Large Language Models (LLMs) have achieved high accuracy on medical question-answer (QA) benchmarks, yet their capacity for flexible clinical reasoning has been debated. Here, we a…
cs.CL2025
Limitations of Large Language Models in Clinical Problem-Solving Arising from Inflexible Reasoning
Jonathan Kim, Anna Podlasek, Kie Shidara +3
Large Language Models (LLMs) have attained human-level accuracy on medical question-answer (QA) benchmarks. However, their limitations in navigating open-ended clinical scenarios h…