5 papers · 1 filter
Fast Byte Latent Transformer
Julie Kallini, Artidoro Pagnoni, Tomasz Limisiewicz +5
Recent byte-level language models (LMs) match the performance of token-level models without relying on subword vocabularies, yet their utility is limited by slow, byte-by-byte auto…
Synthetic Data for any Differentiable Target
Tristan Thrush, Sung Min Park, Herman Brunborg +5
What are the limits of controlling language models via synthetic training data? We develop a reinforcement learning (RL) primitive, the Dataset Policy Gradient (DPG), which can pre…
Improving Pretraining Data Using Perplexity Correlations
Tristan Thrush, Christopher Potts, Tatsunori Hashimoto
Quality pretraining data is often seen as the key to high-performance language models. However, progress in understanding pretraining data has been slow due to the costly pretraini…
I am a Strange Dataset: Metalinguistic Tests for Language Models
Tristan Thrush, Jared Moore, Miguel Monares +2
Statements involving metalinguistic self-reference ("This paper has six sections.") are prevalent in many domains. Can current large language models (LLMs) handle such language? In…
Mission: Impossible Language Models
Julie Kallini, Isabel Papadimitriou, Richard Futrell +2
Chomsky and others have very directly claimed that large language models (LLMs) are equally capable of learning languages that are possible and impossible for humans to learn. Howe…