5 papers · 1 filter
Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind
Christopher Ackerman
The ability to represent oneself and others as agents with knowledge, intentions, and belief states that guide their behavior - Theory of Mind - is a human universal that enables u…
Evidence for Limited Metacognition in LLMs
Christopher Ackerman
The possibility of LLM self-awareness and even sentience is gaining increasing public attention and has major safety and policy implications, but the science of measuring them is s…
Mitigating Many-Shot Jailbreaking
Christopher M. Ackerman, Nina Panickssery
Many-shot jailbreaking (MSJ) is an adversarial technique that exploits the long context windows of modern LLMs to circumvent model safety training by including in the prompt many e…
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
Christopher Ackerman, Nina Panickssery
It has been reported that LLMs can recognize their own writing. As this has potential implications for AI safety, yet is relatively understudied, we investigate the phenomenon, see…
Representation Tuning
Christopher M. Ackerman
Activation engineering is becoming increasingly popular as a means of online control of large language models (LLMs). In this work, we extend the idea of inference-time steering wi…