collaborators

14 papers

cs.CY2026

When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games

Jerick Shi, Terry Jingcheng Zhang, Bernhard Schölkopf +2

As large language models are deployed as autonomous agents that communicate intentions before acting, a critical safety question is whether agents that publicly commit to actions w…

cs.GT2026

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas

Emanuel Tewolde, Xiao Zhang, David Guzman Piedrahita +2

It is increasingly important that LLM agents interact effectively and safely with other goal-pursuing agents, yet, recent works report the opposite trend: LLMs with stronger reason…

cs.GT2026

Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI

Xuanqiang Angelo Huang, Charlie Tharas, Samuele Marro +4

Ensuring that AI agents behave safely and beneficially when interacting with other parties has emerged as one of the central challenges of modern AI safety. While mechanism design,…

cs.LG2026

Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints

Xinge Liu, Terry Jingchen Zhang, Bernhard Schölkopf +2

The rise of autonomous AI agents suggests that dynamic benchmark environments with built-in feedback on scientifically grounded tasks are needed to evaluate the capabilities of the…

cs.CL2026

Evaluating Cooperation in LLM Social Groups through Elected Leadership

Ryan Faulkner, Anushka Deshpande, David Guzman Piedrahita +2

Governing common-pool resources requires agents to develop enduring strategies through cooperation and self-governance to avoid collective failure. While foundation models have sho…

cs.CV2026

When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models

Francesco Ortu, Zhijing Jin, Diego Doimo +1

Vision-language models (VLMs) increasingly combine visual and textual information to perform complex tasks. However, conflicts between their internal knowledge and external visual…