1 citations · 1 across the 5 of their papers we have counts for
3 papers · 1 filter
When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games
Jerick Shi, Terry Jingcheng Zhang, Bernhard Schölkopf +2
As large language models are deployed as autonomous agents that communicate intentions before acting, a critical safety question is whether agents that publicly commit to actions w…
Cheap Talk, Empty Promise: Frontier LLMs easily break public promises for self-interest
Jerick Shi, Terry Jingcheng Zhang, Zhijing Jin +1
Large language models are increasingly deployed as autonomous agents in multi-agent settings where they communicate intentions and take consequential actions with limited human ove…
From Sycophancy to Deception: A Unified Taxonomy for LLM Spontaneous Misalignment
Jerick Shi, Terry Jingcheng Zhang, Zhijing Jin +1
Large language models (LLMs) could produce systematically misaligned output, from hallucinated citations to strategic deception of evaluators, yet these phenomena are studied by se…