4 papers
DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking
Matt L. Wiemann, Lindsay M. Smith, Peter Melchior +4
Frontier LLMs now perform strongly across a wide range of physics evaluations, but it is hard to disentangle genuine reasoning from recall of established science. We introduce Disc…
Large language models and the entropy of English
Colin Scheibner, Lindsay M. Smith, William Bialek
We use large language models (LLMs) to uncover long-ranged structure in English texts from a variety of sources. The conditional entropy or code length in many cases continues to d…
ALICE: An Interpretable Neural Architecture for Generalization in Substitution Ciphers
Jeff Shen, Lindsay M. Smith
We present cryptogram solving as an ideal testbed for studying neural network reasoning and generalization; models must decrypt text encoded with substitution ciphers, choosing fro…
When can in-context learning generalize out of task distribution?
Chase Goddard, Lindsay M. Smith, Vudtiwat Ngampruetikorn +1
In-context learning (ICL) is a remarkable capability of pretrained transformers that allows models to generalize to unseen tasks after seeing only a few examples. We investigate em…