322 citations · 523 across the 4 of their papers we have counts for
4 papers
Mixtral of Experts
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux +23
We introduce Mixtral 8x7B, a Sparse Mixture of Experts (SMoE) language model. Mixtral has the same architecture as Mistral 7B, with the difference that each layer is composed of 8…
Mistral 7B
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch +15
We introduce Mistral 7B v0.1, a 7-billion-parameter language model engineered for superior performance and efficiency. Mistral 7B outperforms Llama 2 13B across all evaluated bench…
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang +51
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…
Training Transformers Together
Alexander Borzunov, Max Ryabinin, Tim Dettmers +5
The infrastructure necessary for training state-of-the-art models is becoming overly expensive, which makes training such models affordable only to large corporations and instituti…