3 papers
cs.AI2025
Evaluation of Causal Reasoning for Large Language Models in Contextualized Clinical Scenarios of Laboratory Test Interpretation
Balu Bhasuran, Mattia Prosperi, Karim Hanna +3
This study evaluates causal reasoning in large language models (LLMs) using 99 clinically grounded laboratory test scenarios aligned with Pearl's Ladder of Causation: association,…
cs.LG2025
HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class
James V. Roggeveen, Erik Y. Wang, Will Flintoft +42
Large language models (LLMs) have shown remarkable progress in mathematical problem-solving, but evaluation has largely focused on problems that have exact analytical solutions or…
cs.CL2024
Evaluating the Impact of Lab Test Results on Large Language Models Generated Differential Diagnoses from Clinical Case Vignettes
Balu Bhasuran, Qiao Jin, Yuzhang Xie +6
Differential diagnosis is crucial for medicine as it helps healthcare providers systematically distinguish between conditions that share similar symptoms. This study assesses the i…