2 papers
cs.CL2026
The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
Christina Lu, Jack Gallagher, Jonathan Michala +2
Large language models can represent a variety of personas but typically default to a helpful Assistant identity cultivated during post-training. We investigate the structure of the…
cs.CL2026
Emergent Introspective Awareness in Large Language Models
Jack Lindsey
We investigate whether large language models can introspect on their internal states. It is difficult to answer this question through conversation alone, as genuine introspection c…