4 papers
Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them
Kevin Zhou, Lisa Alazraki, Kris Cao +1
Pre-training data mixtures are commonly tuned by running small-scale experiments and extrapolating to the target training budget. When high-quality data is scarce and must be repea…
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
Sam Bowyer, Acyr Locatelli, Kris Cao
Efficient benchmarking techniques aim to lower the computational cost of evaluating LLMs by predicting full benchmark scores using only a subset of a benchmark's questions. By refr…
One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers
Diana Abagyan, Alejandro R. Salamanca, Andres Felipe Cruz-Salinas +6
Pretraining massively multilingual Large Language Models (LLMs) for many languages at once is challenging due to limited model capacity, scarce high-quality data, and compute const…
Command A: An Enterprise-Ready Large Language Model
Team Cohere, :, Aakanksha +227
In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command A is an agent-optimised…