4 papers
CooperBench: Why Coding Agents Cannot be Your Teammates Yet
Arpandeep Khatua, Hao Zhu, Peter Tran +8
Resolving team conflicts requires not only task-specific competence, but also social intelligence to find common ground and build consensus. As AI agents increasingly collaborate o…
DECEPTICON: How Dark Patterns Manipulate Web Agents
Phil Cuvin, Hao Zhu, Diyi Yang
Deceptive UI designs, widely instantiated across the web and commonly known as dark patterns, manipulate users into performing actions misaligned with their goals. In this paper, w…
Real-Time Reasoning Agents in Evolving Environments
Yule Wen, Yixin Ye, Yanzhe Zhang +2
Agents in the real world must make not only logical but also timely judgments. This requires continuous awareness of the dynamic environment: hazards emerge, opportunities arise, a…
AutoLibra: Agent Metric Induction from Open-Ended Human Feedback
Hao Zhu, Phil Cuvin, Xinkai Yu +3
Agents are predominantly evaluated and optimized via task success metrics, which are coarse, rely on manual design from experts, and fail to reward intermediate emergent behaviors.…