4 papers
Interpretability Can Be Actionable
Hadas Orgad, Fazl Barez, Tal Haklay +9
Interpretability aims to explain the behavior of deep neural networks. Despite rapid growth, there is mounting concern that much of this work has not translated into practical impa…
Hidden Breakthroughs in Language Model Training
Sara Kangaslahti, Elan Rosenfeld, Naomi Saphra
Loss curves are smooth during most of model training, so visible discontinuities stand out as possible conceptual breakthroughs. Studying these breakthroughs enables a deeper under…
ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in Context
Victoria R. Li, Yida Chen, Naomi Saphra
While the biases of language models in production are extensively documented, the biases of their guardrails have been neglected. This paper studies how contextual information abou…
A Taxonomy of Transcendence
Natalie Abreu, Edwin Zhang, Eran Malach +1
Although language models are trained to mimic humans, the resulting systems display capabilities beyond the scope of any one person. To understand this phenomenon, we use a control…