5 papers
Curriculum-Guided Layer Scaling for Language Model Pretraining
Karanpartap Singh, Neil Band, Ehsan Adeli
As the cost of pretraining large language models grows, there is continued interest in strategies to improve learning efficiency during this core training stage. Motivated by cogni…
Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
Benjamin Rozonoyer, Jacopo Minniti, Dhruvesh Patel +4
When Masked Diffusion Models (MDMs) generate sequences through iterative refinement, the rich internal computation over masked positions is discarded, forcing every subsequent refi…
Synthetic Data for any Differentiable Target
Tristan Thrush, Sung Min Park, Herman Brunborg +5
What are the limits of controlling language models via synthetic training data? We develop a reinforcement learning (RL) primitive, the Dataset Policy Gradient (DPG), which can pre…
Reasoning to Learn from Latent Thoughts
Yangjun Ruan, Neil Band, Chris J. Maddison +1
Compute scaling for language model (LM) pretraining has outpaced the growth of human-written texts, leading to concerns that data will become the bottleneck to LM scaling. To conti…
Synthetic continued pretraining
Zitong Yang, Neil Band, Shuangping Li +2
Pretraining on large-scale, unstructured internet text enables language models to acquire a significant amount of world knowledge. However, this knowledge acquisition is data-ineff…