5 papers
When Role-playing, Do Models Believe What They Say?
Benjamin Sturgeon, David Africa, Sid Black
Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite. Recent work argues that persona adoption is fundamental to how lang…
More Capable, Less Cooperative? When LLMs Fail At Zero-Cost Collaboration
Advait Yadav, Sid Black, Oliver Sourbut
Large language model (LLM) agents increasingly coordinate in multi-agent systems, yet we lack an understanding of where and why cooperation fails. Many real-world coordination prob…
Do Large Language Models Know What They Are Capable Of?
Casey O. Barkan, Sid Black, Oliver Sourbut
We investigate whether large language models (LLMs) can predict whether they will succeed on a given task and whether their predictions improve as they progress through multi-step…
Auditing Games for Sandbagging
Jordan Taylor, Sid Black, Dillon Bowen +10
Future AI systems could conceal their capabilities ('sandbagging') during evaluations, potentially misleading developers and auditors. We stress-tested sandbagging detection techni…
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
Sid Black, Asa Cooper Stickland, Jake Pencharz +7
Uncontrollable autonomous replication of language model agents poses a critical safety risk. To better understand this risk, we introduce RepliBench, a suite of evaluations designe…