2 papers
cs.CL2026
Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models
Kalyani Marathe, Artidoro Pagnoni, Tomasz Limisiewicz +4
Small models are made more capable through distillation from a larger one that shares their tokenization scheme. However, do distilled byte and token models behave similarly in ter…
cs.LG2024
DataComp-LM: In search of the next generation of training sets for language models
Jeffrey Li, Alex Fang, Georgios Smyrnis +56
We introduce DataComp for Language Models (DCLM), a testbed for controlled dataset experiments with the goal of improving language models. As part of DCLM, we provide a standardize…