Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Training Language Models to Explain Their Own Computations
Belinda Z. Li, Zifan Carl Guo, Vincent Huang +2
Can language models (LMs) learn to faithfully describe their internal computations? Are they better able to describe themselves than other models? We study the extent to which LMs'…
cs.CL2025
(How) Do Language Models Track State?
Belinda Z. Li, Zifan Carl Guo, Jacob Andreas
Transformer language models (LMs) exhibit behaviors -- from storytelling to code generation -- that seem to require tracking the unobserved state of an evolving world. How do they…