collaborators

5 papers

cs.CL2026

Spine-Branch Coordination for Multi-agent Computer Use

Mian Zhang, Manasi Sharma, Sheng Zhang +5

Computer use agents (CUAs) are increasingly deployed as multi-agent systems that decompose a task into multiple subtasks executed across parallel virtual machines (VMs). However, a…

cs.LG2026

Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

Minglai Yang, Xinyu Guo, Utkarsh Tyagi +6

Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer. The…

cs.AI2026

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

Veronica Chatrath, Bryan Zhu, George Pu +16

Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, lo…

cs.CY2026

The Limits of AI Data Transparency Policy: Three Disclosure Fallacies

Judy Hanwen Shen, Ken Liu, Angelina Wang +7

Data transparency has emerged as a rallying cry for addressing concerns about AI: data quality, privacy, and copyright chief among them. Yet while these calls are crucial for accou…

cs.CY2025

The California Report on Frontier AI Policy

Rishi Bommasani, Scott R. Singer, Ruth E. Appel +20

The innovations emerging at the frontier of artificial intelligence (AI) are poised to create historic opportunities for humanity but also raise complex policy challenges. Continue…