132 citations · 242 across the 6 of their papers we have counts for
4 papers · 1 filter
ConstitutionalExperts: Training a Mixture of Principle-based Prompts
Savvas Petridis, Ben Wedin, Ann Yuan +2
Large language models (LLMs) are highly capable at a variety of tasks given the right prompt, but writing one is still a difficult and tedious process. In this work, we introduce C…
Wordcraft: a Human-AI Collaborative Editor for Story Writing
Andy Coenen, Luke Davis, Daphne Ippolito +2
As neural language models grow in effectiveness, they are increasingly being applied in real-world settings. However these applications tend to be limited in the modes of interacti…
An Interpretability Illusion for BERT
Tolga Bolukbasi, Adam Pearce, Ann Yuan +4
We describe an "interpretability illusion" that arises when analyzing the BERT model. Activations of individual neurons in the network may spuriously appear to encode a single, sim…
The Language Interpretability Tool: Extensible, Interactive Visualizations and Analysis for NLP Models
Ian Tenney, James Wexler, Jasmijn Bastings +8
We present the Language Interpretability Tool (LIT), an open-source platform for visualization and understanding of NLP models. We focus on core questions about model behavior: Why…