Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs
Rohith Reddy Bellibatlu, Manpreet Singh, Zhoutian Han +1
A clinical agent benchmark can report the same verdict on identical inputs while the agent files a materially different order on each run. Such agents order tests, request medicati…
cs.CL2026
Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor
Rohith Reddy Bellibatlu, Manpreet Singh, Deepak Parashar +1
Counterfactual audits are the standard tool for checking whether a clinical agent treats demographically distinct but clinically identical patients differently. They report a flip…