4 papers
Priors in Time: Missing Inductive Biases for Language Model Interpretability
Ekdeep Singh Lubana, Can Rager, Sai Sumedh R. Hindupur +13
Recovering meaningful concepts from language model activations is a central aim of interpretability. While existing feature extraction methods aim to identify concepts that are ind…
Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
Alexa R. Tartaglini, Satchel Grant, Daniel Wurgaft +2
Data visualizations are vital components of many scientific articles and news stories. Current vision-language models (VLMs) still struggle on basic data visualization understandin…
In-Context Learning Strategies Emerge Rationally
Daniel Wurgaft, Ekdeep Singh Lubana, Core Francisco Park +3
Recent work analyzing in-context learning (ICL) has identified a broad set of strategies that describe model behavior in different experimental conditions. We aim to unify these fi…
Scaling up the think-aloud method
Daniel Wurgaft, Ben Prystawski, Kanishk Gandhi +3
The think-aloud method, where participants voice their thoughts as they solve a task, is a valuable source of rich data about human reasoning processes. Yet, it has declined in pop…