4 papers
What Does it Mean for a Neural Network to Learn a "World Model"?
Kenneth Li, Fernanda Viégas, Martin Wattenberg
We propose a set of precise criteria for saying a neural net learns and uses a "world model." The goal is to give an operational meaning to terms that are often used informally, in…
When Bad Data Leads to Good Models
Kenneth Li, Yida Chen, Fernanda Viégas +1
In large language model (LLM) pretraining, data quality is believed to determine model quality. In this paper, we re-examine the notion of "quality" from the perspective of pre- an…
Communicating Activations Between Language Model Agents
Vignav Ramesh, Kenneth Li
Communication between multiple language model (LM) agents has been shown to scale up the reasoning ability of LMs. While natural language has been the dominant medium for inter-LM…
Designing a Dashboard for Transparency and Control of Conversational AI
Yida Chen, Aoyu Wu, Trevor DePodesta +9
Conversational LLMs function as black box systems, leaving users guessing about why they see the output they do. This lack of transparency is potentially problematic, especially gi…