Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
Veronica Chatrath, Bryan Zhu, George Pu +16
Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, lo…
cs.AI2026
ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks
Vincent Siu, Manasi Sharma, Dawn Song +3
Computer use agents are evaluated almost exclusively on atomic desktop tasks, but realistic desktop work requires sustaining state across multiple objectives. We study this gap wit…