3 papers
cs.AI2026
When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models
Yujun Zhou, Christopher M. Ackerman
Recent work on preference elicitation in large language models (LLMs) has demonstrated that, when given a series of choices between two outcomes, LLMs reveal a coherent, model-spec…
cs.LG2025
Mitigating Many-Shot Jailbreaking
Christopher M. Ackerman, Nina Panickssery
Many-shot jailbreaking (MSJ) is an adversarial technique that exploits the long context windows of modern LLMs to circumvent model safety training by including in the prompt many e…
cs.LG2024
Representation Tuning
Christopher M. Ackerman
Activation engineering is becoming increasingly popular as a means of online control of large language models (LLMs). In this work, we extend the idea of inference-time steering wi…