2 papers
cs.AI2026
Persona Cartography: Charting Language Model Personality Traits in Weight Space
Luke Baines, Anton Gonzalvez Hawthorne, Mariia Koroliuk +4
Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tools for decomposing, measuring, and control…
cs.CL2025
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
Iván Vicente Moreno Cencerrado, Arnau Padrés Masdemont, Anton Gonzalvez Hawthorne +2
Do large language models (LLMs) anticipate when they will answer correctly? To study this, we extract activations after a question is read but before any tokens are generated, and…