activity
20242026
most citedSynthetic continued pretraining

1 citations · 1 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CL2026

Synthetic Data for any Differentiable Target

Tristan Thrush, Sung Min Park, Herman Brunborg +5

What are the limits of controlling language models via synthetic training data? We develop a reinforcement learning (RL) primitive, the Dataset Policy Gradient (DPG), which can pre…

cs.LG2026

Data-efficient pre-training by scaling synthetic megadocs

Konwoo Kim, Suhas Kotha, Yejin Choi +3

Synthetic data augmentation has emerged as a promising solution when pre-training is constrained by data rather than compute. We study how to design synthetic data algorithms that…

cs.CL2026

Agentic Adversarial QA for Improving Domain-Specific LLMs

Vincent Grari, Ciprian Tomoiaga, Sylvain Lamprier +2

Large Language Models (LLMs), despite extensive pretraining on broad internet corpora, often struggle to adapt effectively to specialized domains. There is growing interest in fine…

cs.CL2025

Synthetic bootstrapped pretraining

Zitong Yang, Aonan Zhang, Hong Liu +4

We introduce Synthetic Bootstrapped Pretraining (SBP), a language model (LM) pretraining procedure that first learns a model of relations between documents from the pretraining dat…

cs.LG2025

Reasoning to Learn from Latent Thoughts

Yangjun Ruan, Neil Band, Chris J. Maddison +1

Compute scaling for language model (LM) pretraining has outpaced the growth of human-written texts, leading to concerns that data will become the bottleneck to LM scaling. To conti…

cs.LG20241 cited

Synthetic continued pretraining

Zitong Yang, Neil Band, Shuangping Li +2

Pretraining on large-scale, unstructured internet text enables language models to acquire a significant amount of world knowledge. However, this knowledge acquisition is data-ineff…