751 citations · 1k across the 9 of their papers we have counts for
5 papers · 1 filter
Linearity of Relation Decoding in Transformer Language Models
Evan Hernandez, Arnab Sen Sharma, Tal Haklay +5
Much of the knowledge encoded in transformer language models (LMs) may be expressed in terms of relations: relations between words and their synonyms, entities and their attributes…
Beyond Surface Statistics: Scene Representations in a Latent Diffusion Model
Yida Chen, Fernanda Viégas, Martin Wattenberg
Latent diffusion models (LDMs) exhibit an impressive ability to produce realistic images, yet the inner workings of these models remain mysterious. Even when trained purely on imag…
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
Kenneth Li, Oam Patel, Fernanda Viégas +2
We introduce Inference-Time Intervention (ITI), a technique designed to enhance the "truthfulness" of large language models (LLMs). ITI operates by shifting model activations durin…
The System Model and the User Model: Exploring AI Dashboard Design
Fernanda Viégas, Martin Wattenberg
This is a speculative essay on interface design and artificial intelligence. Recently there has been a surge of attention to chatbots based on large language models, including wide…
AttentionViz: A Global View of Transformer Attention
Catherine Yeh, Yida Chen, Aoyu Wu +3
Transformer models are revolutionizing machine learning, but their inner workings remain mysterious. In this work, we present a new visualization technique designed to help researc…