3 papers
cs.LG2026
From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language
Zachary Wojtowicz, Ayush Nayak, Jacob Andreas
The growing use of statistical learning algorithms to infer human preferences from high-dimensional choice data runs up against a fundamental challenge: choice alternatives typical…
cs.CL2026
Training Language Models to Explain Their Own Computations
Belinda Z. Li, Zifan Carl Guo, Vincent Huang +2
Can language models (LMs) learn to faithfully describe their internal computations? Are they better able to describe themselves than other models? We study the extent to which LMs'…
cs.CL2025
(How) Do Language Models Track State?
Belinda Z. Li, Zifan Carl Guo, Jacob Andreas
Transformer language models (LMs) exhibit behaviors -- from storytelling to code generation -- that seem to require tracking the unobserved state of an evolving world. How do they…