1.1k citations · 1.7k across the 55 of their papers we have counts for
3 papers · 1 filter
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
Feiyang Kang, Newsha Ardalani, Michael Kuchnik +7
Training data plays a crucial role in Large Language Models (LLM) scaling, yet high quality data is of limited supply. Synthetic data techniques offer a potential path toward sides…
Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning
Lili Yu, Bowen Shi, Ramakanth Pasunuru +24
We present CM3Leon (pronounced "Chameleon"), a retrieval-augmented, token-based, decoder-only multi-modal language model capable of generating and infilling both text and images. C…
Supporting Clustering with Contrastive Learning
Dejiao Zhang, Feng Nan, Xiaokai Wei +6
Unsupervised clustering aims at discovering the semantic categories of data according to some distance measured in the representation space. However, different categories often ove…