1 paper · 1 filter
Raphaël Sarfati, Eric Bigelow, Daniel Wurgaft +6
Large language models (LLMs) form implicit beliefs (posteriors over latent variables) from prompts, but we lack a mechanistic account of how these beliefs are encoded in representa…