4 papers
Pitfalls in Evaluating Interpretability Agents
Tal Haklay, Nikhil Prakash, Sana Pandey +5
Automated interpretability systems aim to reduce the need for human labor and scale analysis to increasingly large models and diverse tasks. Recent efforts toward this goal leverag…
Modeling Boundedly Rational Agents with Latent Inference Budgets
Athul Paul Jacob, Abhishek Gupta, Jacob Andreas
We study the problem of modeling a population of agents pursuing unknown goals subject to unknown computational constraints. In standard models of bounded rationality, sub-optimal…
Regularized Conventions: Equilibrium Computation as a Model of Pragmatic Reasoning
Athul Paul Jacob, Gabriele Farina, Jacob Andreas
We present a model of pragmatic language understanding, where utterances are produced and understood by searching for regularized equilibria of signaling games. In this model (whic…
The Consensus Game: Language Model Generation via Equilibrium Search
Athul Paul Jacob, Yikang Shen, Gabriele Farina +1
When applied to question answering and other text generation tasks, language models (LMs) may be queried generatively (by sampling answers from their output distribution) or discri…