1 citations · 1 across the 6 of their papers we have counts for
4 papers · 1 filter
What do Language Models Learn and When? The Implicit Curriculum Hypothesis
Emmy Liu, Kaiser Sun, Millicent Li +4
Large language models (LLMs) can perform remarkably complex tasks, yet the fine-grained details of how these capabilities emerge during pretraining remain poorly understood. Scalin…
To Memorize or to Retrieve: Scaling the Interaction Between Pretraining and Retrieval
Karan Singh, Michael Yu, Varun Gangal +4
Retrieval-augmented generation (RAG) improves language model (LM) performance by providing relevant context at test time for knowledge-intensive situations. In this work, we system…
Midtraining Bridges Pretraining and Posttraining Distributions
Emmy Liu, Graham Neubig, Chenyan Xiong
Midtraining, the practice of mixing specialized data with more general pretraining data in an intermediate training phase, has become widespread in language model development, yet…
An Incomplete Loop: Deductive, Inductive, and Abductive Reasoning in Language Models
Emmy Liu, Graham Neubig, Jacob Andreas
Modern language models (LMs) can learn to perform new tasks in different ways: in instruction following, the target task is described explicitly in natural language; in few-shot pr…