5 papers · 1 filter
World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments
Ananya Mantravadi, Harshit Rajgarhia, Prasanna Desikan +1
Clinical protocol-execution tasks -- checking a lab value, applying a threshold, placing a correctly structured FHIR order -- are natural candidates for RL from world feedback: onc…
LLM Doesn't Know What It Doesn't Know: Detecting Epistemic Blind Spots via Cross-Model Attribution Divergence on Clinical Tabular Data
Akshat Dasula, Prasanna Desikan, Jaideep Srivastava
Large language models (LLMs) are increasingly applied to structured clinical data, yet whether they can recognize the limits of their own knowledge on such tasks remains unexplored…
CareTransition-Audit: A Benchmark to Audit Discharge Summaries for Efficient Care Transitions
Akshat Dasula, Prasanna Desikan, Jaideep Srivastava +2
Incomplete or inconsistent discharge documentation drives care fragmentation and avoidable readmissions. Despite its critical role in patient safety, auditing discharge summaries r…
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare
Prasanna Desikan, Harshit Rajgarhia, Shivali Dalmia +1
AI models are increasingly deployed in live clinical environments where they must perform reliably across complex, high-stakes workflows that standard training and validation datas…
Systematic Evaluation of Large Language Models for Post-Discharge Clinical Action Extraction
Shivali Dalmia, Ananya Mantravadi, Prasanna Desikan
The work in this paper evaluates zero-shot and few-shot large language models (LLMs) for safety-critical clinical action extraction using the CLIP discharge-note dataset, with part…