4 papers
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
Veronica Chatrath, Bryan Zhu, George Pu +16
Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, lo…
ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks
Vincent Siu, Manasi Sharma, Dawn Song +3
Computer use agents are evaluated almost exclusively on atomic desktop tasks, but realistic desktop work requires sustaining state across multiple objectives. We study this gap wit…
The Limits of AI Data Transparency Policy: Three Disclosure Fallacies
Judy Hanwen Shen, Ken Liu, Angelina Wang +7
Data transparency has emerged as a rallying cry for addressing concerns about AI: data quality, privacy, and copyright chief among them. Yet while these calls are crucial for accou…
The California Report on Frontier AI Policy
Rishi Bommasani, Scott R. Singer, Ruth E. Appel +20
The innovations emerging at the frontier of artificial intelligence (AI) are poised to create historic opportunities for humanity but also raise complex policy challenges. Continue…