3 papers
cs.LG2026
CooperBench: Why Coding Agents Cannot be Your Teammates Yet
Arpandeep Khatua, Hao Zhu, Peter Tran +8
Resolving team conflicts requires not only task-specific competence, but also social intelligence to find common ground and build consensus. As AI agents increasingly collaborate o…
cs.AI2025
AutoLibra: Agent Metric Induction from Open-Ended Human Feedback
Hao Zhu, Phil Cuvin, Xinkai Yu +3
Agents are predominantly evaluated and optimized via task success metrics, which are coarse, rely on manual design from experts, and fail to reward intermediate emergent behaviors.…
cs.CL2024
DiverseDialogue: A Methodology for Designing Chatbots with Human-Like Diversity
Xiaoyu Lin, Xinkai Yu, Ankit Aich +2
Large Language Models (LLMs), which simulate human users, are frequently employed to evaluate chatbots in applications such as tutoring and customer service. Effective evaluation n…