887 citations · 928 across the 20 of their papers we have counts for
4 papers · 1 filter
Explaining Attention with Program Synthesis
Amiri Hayes, Belinda Z Li, Jacob Andreas
A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic descriptions. In this paper, we propose an ap…
Self-CTRL: Self-Consistency Training with Reinforcement Learning
Itamar Pres, Laura Ruis, Melat Ghebreselassie +2
Language models (LMs) that faithfully describe their own behavior can more easily be audited, understood, and trusted by users. This paper describes Self-Consistency Training with…
LaMPP: Language Models as Probabilistic Priors for Perception and Action
Belinda Z. Li, William Chen, Pratyusha Sharma +1
Language models trained on large text corpora encode rich distributional information about real-world environments and action sequences. This information plays a crucial role in cu…
Linformer: Self-Attention with Linear Complexity
Sinong Wang, Belinda Z. Li, Madian Khabsa +2
Large transformer models have shown extraordinary success in achieving state-of-the-art results in many natural language processing applications. However, training and deploying th…