24 citations · 159 across the 20 of their papers we have counts for
9 papers · 1 filter
FIND: A Function Description Benchmark for Evaluating Interpretability Methods
Sarah Schwettmann, Tamar Rott Shaham, Joanna Materzynska +5
Labeling neural network submodules with human-legible descriptions is useful for many downstream tasks: such descriptions can surface failures, guide interventions, and perhaps eve…
Linearity of Relation Decoding in Transformer Language Models
Evan Hernandez, Arnab Sen Sharma, Tal Haklay +5
Much of the knowledge encoded in transformer language models (LMs) may be expressed in terms of relations: relations between words and their synonyms, entities and their attributes…
Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks
Zhaofeng Wu, Linlu Qiu, Alexis Ross +6
The impressive performance of recent language models across a wide range of tasks suggests that they possess a degree of abstract reasoning skills. Are these skills general and tra…
From Word Models to World Models: Translating from Natural Language to the Probabilistic Language of Thought
Lionel Wong, Gabriel Grand, Alexander K. Lew +4
How does language inform our downstream thinking? In particular, how do humans make meaning from language--and how can we leverage a theory of linguistic meaning to build machines…
The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks
Ziqian Zhong, Ziming Liu, Max Tegmark +1
Do neural networks, trained on well-understood algorithmic tasks, reliably rediscover known algorithms for solving those tasks? Several recent studies, on tasks ranging from group…
Decision-Oriented Dialogue for Human-AI Collaboration
Jessy Lin, Nicholas Tomlin, Jacob Andreas +1
We describe a class of tasks called decision-oriented dialogues, in which AI assistants such as large language models (LMs) must collaborate with one or more humans via natural lan…