5 papers
When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models
Yujun Zhou, Christopher M. Ackerman
Recent work on preference elicitation in large language models (LLMs) has demonstrated that, when given a series of choices between two outcomes, LLMs reveal a coherent, model-spec…
Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind
Christopher Ackerman
The ability to represent oneself and others as agents with knowledge, intentions, and belief states that guide their behavior - Theory of Mind - is a human universal that enables u…
Evidence for Limited Metacognition in LLMs
Christopher Ackerman
The possibility of LLM self-awareness and even sentience is gaining increasing public attention and has major safety and policy implications, but the science of measuring them is s…
Mitigating Many-Shot Jailbreaking
Christopher M. Ackerman, Nina Panickssery
Many-shot jailbreaking (MSJ) is an adversarial technique that exploits the long context windows of modern LLMs to circumvent model safety training by including in the prompt many e…
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
Christopher Ackerman, Nina Panickssery
It has been reported that LLMs can recognize their own writing. As this has potential implications for AI safety, yet is relatively understudied, we investigate the phenomenon, see…