5 papers
Subliminal Steering: Stronger Encoding of Hidden Signals
George Morgulis, John Hewitt
Subliminal learning describes a student language model inheriting a behavioral bias by fine-tuning on seemingly innocuous data generated by a biased teacher model. Prior work has b…
Improving Parametric Knowledge Access in Reasoning Language Models
Melody Ma, John Hewitt
We study reasoning for accessing world knowledge stored in a language model's parameters. For example, recalling that Canberra is Australia's capital may benefit from thinking thro…
Neologism Learning for Controllability and Self-Verbalization
John Hewitt, Oyvind Tafjord, Robert Geirhos +1
Humans invent new words when there is a rising demand for a new useful concept (e.g., doomscrolling). We explore and validate a similar idea in our communication with LLMs: introdu…
Because we have LLMs, we Can and Should Pursue Agentic Interpretability
Been Kim, John Hewitt, Neel Nanda +2
The era of Large Language Models (LLMs) presents a new opportunity for interpretability--agentic interpretability: a multi-turn conversation with an LLM wherein the LLM proactively…
We Can't Understand AI Using our Existing Vocabulary
John Hewitt, Robert Geirhos, Been Kim
This position paper argues that, in order to understand AI, we cannot rely on our existing vocabulary of human words. Instead, we should strive to develop neologisms: new words tha…