4 papers
Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor
Rohith Reddy Bellibatlu, Manpreet Singh, Deepak Parashar +1
Counterfactual audits are the standard tool for checking whether a clinical agent treats demographically distinct but clinically identical patients differently. They report a flip…
ExecuGraph: A Multi-Agent, Execution-Grounded Framework for Reliable Backend Code Synthesis with Large Language Models
Sai Deekshith Lekkala, Jothi Prabha Appadurai, Rohith Reddy Bellibatlu +1
Large Language Models generate plausible backend code, but a single-pass paradigm provides no guarantee of correctness or runtime reliability. We present ExecuGraph, a multi-agent…
RISED: A Pre-Deployment Evaluation Framework for High-Stakes AI Decision-Support Systems, with Application to Healthcare
Rohith Reddy Bellibatlu, Manpreet Singh, Yash Jajoo +2
Clinical decision-support systems are expert systems whose recommendations clinicians act on directly, yet they are usually cleared on one aggregate accuracy number from a held-out…
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
Rohith Reddy Bellibatlu, Edward Raff, Wenbin Zhang
Large language models are widely adopted as automated evaluation judges, yet the stability of their verdicts under semantically equivalent prompt rephrasings remains largely unexam…