4 papers
Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
Vincent Huang, Dami Choi, Daniel D. Johnson +2
Interpreting the internal activations of neural networks can produce more faithful explanations of their behavior, but is difficult due to the complex structure of activation space…
Eliciting Language Model Behaviors with Investigator Agents
Xiang Lisa Li, Neil Chowdhury, Daniel D. Johnson +4
Language models exhibit complex, diverse behaviors when prompted with free-form text, making it difficult to characterize the space of possible outputs. We study the problem of beh…
Penzai + Treescope: A Toolkit for Interpreting, Visualizing, and Editing Models As Data
Daniel D. Johnson
Much of today's machine learning research involves interpreting, modifying or visualizing models after they are trained. I present Penzai, a neural network library designed to simp…
A density estimation perspective on learning from pairwise human preferences
Vincent Dumoulin, Daniel D. Johnson, Pablo Samuel Castro +2
Learning from human feedback (LHF) -- and in particular learning from pairwise preferences -- has recently become a crucial ingredient in training large language models (LLMs), and…