7 papers
Probing the Misaligned Thinking Process of Language Models
Kaiwen Zhou, Constantin Venhoff, Jonathan Michala +2
Large language models exhibit a growing range of misaligned behaviors such as strategic deception, sandbagging, and self-preservation. As they are increasingly deployed in high-sta…
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
Xiaoyuan Zhu, Yaowen Ye, Tianyi Qiu +6
As API access becomes a primary interface to large language models (LLMs), users often interact with black-box systems that offer little transparency into the deployed model. To re…
Three Concrete Challenges and Two Hopes for the Safety of Unsupervised Elicitation
Callum Canavan, Aditya Shrivastava, Allison Qi +2
To steer language models towards truthful outputs on tasks which are beyond human capability, previous work has suggested training models on easy tasks to steer them on harder ones…
Abstractive Red-Teaming of Language Model Character
Nate Rahn, Allison Qi, Avery Griffin +3
We want language model assistants to conform to a character specification, which asserts how the model should act across diverse user interactions. While models typically follow th…
The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
Christina Lu, Jack Gallagher, Jonathan Michala +2
Large language models can represent a variety of personas but typically default to a helpful Assistant identity cultivated during post-training. We investigate the structure of the…
Refined Elementary Capacities from Symplectic Field Theory
Jonathan Michala
We extend the family of capacities given by McDuff and Siegel by including a constraint on the number of positive asymptotically cylindrical ends of curves showing up in the…