18 papers
Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments
Keyu He, Xuhui Zhou, Maarten Sap
LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and improving these social skills is…
OdysSim: Building Foundation Models for Human Behavior Simulation
Xuhui Zhou, Weiwei Sun, Weihua Du +6
Large language models are increasingly deployed as human simulators for interactive evaluation and social simulation. Yet helpfulness-driven post-training pulls them toward a homog…
Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios
Giuseppe Attanasio, Beatrice Savoldi, Daniel Chechelnitsky +4
Speech translation (ST) is increasingly adopted in user applications, yet its evaluation largely focuses on decontextualized testbeds and holistic quality, rather than end users' c…
SOTOPIA-TOM: Evaluating Information Management in Multi-Agent Interaction with Theory of Mind
Yashwanth YS, Ruichen Wang, Shihua Zeng +4
As LLM-based agents are increasingly interacting in multi-party settings, they need to properly handle information asymmetry, i.e., knowing when and to whom to disclose information…
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
Mingqian Zheng, Malia Morgan, Liwei Jiang +2
Current LLM safety alignment techniques improve model robustness against adversarial attacks, but overlook whether and how LLMs can recover helpfulness when benign users clarify th…
GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses
Jimin Mun, Chani Jung, Xuhui Zhou +2
While LLMs hold significant potential to transform scientific research, we advocate for their use to augment and empower researchers rather than to automate research without human…