1 paper
Kie Shidara, Preethi Prem, Jonathan Kim +4
Large Language Models (LLMs) have achieved high accuracy on medical question-answer (QA) benchmarks, yet their capacity for flexible clinical reasoning has been debated. Here, we a…