collaborators

7 papers

cs.AI2026

Probing the Misaligned Thinking Process of Language Models

Kaiwen Zhou, Constantin Venhoff, Jonathan Michala +2

Large language models exhibit a growing range of misaligned behaviors such as strategic deception, sandbagging, and self-preservation. As they are increasingly deployed in high-sta…

cs.CR2026

Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test

Xiaoyuan Zhu, Yaowen Ye, Tianyi Qiu +6

As API access becomes a primary interface to large language models (LLMs), users often interact with black-box systems that offer little transparency into the deployed model. To re…

cs.LG2026

Three Concrete Challenges and Two Hopes for the Safety of Unsupervised Elicitation

Callum Canavan, Aditya Shrivastava, Allison Qi +2

To steer language models towards truthful outputs on tasks which are beyond human capability, previous work has suggested training models on easy tasks to steer them on harder ones…

cs.LG2026

Abstractive Red-Teaming of Language Model Character

Nate Rahn, Allison Qi, Avery Griffin +3

We want language model assistants to conform to a character specification, which asserts how the model should act across diverse user interactions. While models typically follow th…

cs.CL2026

The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models

Christina Lu, Jack Gallagher, Jonathan Michala +2

Large language models can represent a variety of personas but typically default to a helpful Assistant identity cultivated during post-training. We investigate the structure of the…

math.SG2025

Refined Elementary Capacities from Symplectic Field Theory

Jonathan Michala

We extend the family of capacities given by McDuff and Siegel by including a constraint on the number of positive asymptotically cylindrical ends of curves showing up in the…