136 citations · 136 across the 2 of their papers we have counts for
2 papers
cs.CL2024
Nemotron-4 15B Technical Report
Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings +24
We introduce Nemotron-4 15B, a 15-billion-parameter large multilingual language model trained on 8 trillion text tokens. Nemotron-4 15B demonstrates strong performance when assesse…
cs.LG2024★ 136 cited
Mixtral of Experts
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux +23
We introduce Mixtral 8x7B, a Sparse Mixture of Experts (SMoE) language model. Mixtral has the same architecture as Mistral 7B, with the difference that each layer is composed of 8…