2 papers
cs.CL2026
Probing Persona-Dependent Preferences in Language Models
Oscar Gilg, Pierre Beckmann, Daniel Paleka +1
Large language models (LLMs) can be said to have preferences: they reliably pick certain tasks and outputs over others, and preferences shaped by post-training and system prompts a…
cs.AI2026
Split Personality Training: Revealing Latent Knowledge Through Alternate Personalities
Florian Dietz, William Wale, Oscar Gilg +5
Detecting misalignment in large language models is challenging because models may learn to conceal misbehavior during training. Standard auditing techniques fall short: black-box m…