Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
READY or Not: Reliable Enterprise Agent Deployment
Veronica Chatrath, Bryan Zhu, Jingxuan Fan +15
An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent benchmarks measure whether an agent can complete realistic professional work, w…
cs.AI2026
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
Veronica Chatrath, Bryan Zhu, George Pu +16
Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, lo…