136 citations · 141 across the 3 of their papers we have counts for
3 papers
Magistral
Mistral-AI, :, Abhinav Rastogi +98
We introduce Magistral, Mistral's first reasoning model and our own scalable reinforcement learning (RL) pipeline. Instead of relying on existing implementations and RL traces dist…
Scaling Laws for Fine-Grained Mixture of Experts
Jakub Krajewski, Jan Ludziejewski, Kamil Adamczewski +9
Mixture of Experts (MoE) models have emerged as a primary solution for reducing the computational cost of Large Language Models. In this work, we analyze their scaling properties,…
Mixtral of Experts
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux +23
We introduce Mixtral 8x7B, a Sparse Mixture of Experts (SMoE) language model. Mixtral has the same architecture as Mistral 7B, with the difference that each layer is composed of 8…