7 papers
World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments
Ananya Mantravadi, Harshit Rajgarhia, Prasanna Desikan +1
Clinical protocol-execution tasks -- checking a lab value, applying a threshold, placing a correctly structured FHIR order -- are natural candidates for RL from world feedback: onc…
Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows
Karthikeya Aditya Vissa, Sankalp Mane, Ananya Mantravadi +2
Large language models are trained to predict the next token, not to act inside a specific API. In niche enterprise SaaS workflows -- where success means hitting the right endpoint…
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare
Prasanna Desikan, Harshit Rajgarhia, Shivali Dalmia +1
AI models are increasingly deployed in live clinical environments where they must perform reliably across complex, high-stakes workflows that standard training and validation datas…
Systematic Evaluation of Large Language Models for Post-Discharge Clinical Action Extraction
Shivali Dalmia, Ananya Mantravadi, Prasanna Desikan
The work in this paper evaluates zero-shot and few-shot large language models (LLMs) for safety-critical clinical action extraction using the CLIP discharge-note dataset, with part…
ART: Action-based Reasoning Task Benchmarking for Medical AI Agents
Ananya Mantravadi, Shivali Dalmia, Abhishek Mukherji
Reliable clinical decision support requires medical AI agents capable of safe, multi-step reasoning over structured electronic health records (EHRs). While large language models (L…
Scaling Success: A Systematic Review of Peer Grading Strategies for Accuracy, Efficiency, and Learning in Contemporary Education
Uchswas Paul, Ananya Mantravadi, Jash Shah +4
Peer grading has emerged as a scalable solution for assessment in large and online classrooms, offering both logistical efficiency and pedagogical value. However, designing effecti…