Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Tracing Persona Vectors Through LLM Pretraining
Viktor Moskvoretskii, Dominik Glandorf, Jorge Medina Moreira +2
How large language models internally represent high-level behaviors is a core interpretability question with direct relevance to AI safety: it determines what we can detect, audit,…
cs.CL2025
Controllable Context Sensitivity and the Knob Behind It
Julian Minder, Kevin Du, Niklas Stoehr +4
When making predictions, a language model must trade off how much it relies on its context vs. its prior knowledge. Choosing how sensitive the model is to its context is a fundamen…
cs.CL2024
Activation Scaling for Steering and Interpreting Language Models
Niklas Stoehr, Kevin Du, Vésteinn Snæbjarnarson +3
Given the prompt "Rome is in", can we steer a language model to flip its prediction of an incorrect token "France" to a correct token "Italy" by only multiplying a few relevant act…