322 citations · 467 across the 3 of their papers we have counts for
3 papers
Mixtral of Experts
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux +23
We introduce Mixtral 8x7B, a Sparse Mixture of Experts (SMoE) language model. Mixtral has the same architecture as Mistral 7B, with the difference that each layer is composed of 8…
Mistral 7B
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch +15
We introduce Mistral 7B v0.1, a 7-billion-parameter language model engineered for superior performance and efficiency. Mistral 7B outperforms Llama 2 13B across all evaluated bench…
Three ways to improve feature alignment for open vocabulary detection
Relja Arandjelović, Alex Andonian, Arthur Mensch +3
The core problem in zero-shot open vocabulary detection is how to align visual and text features, so that the detector performs well on unseen classes. Previous approaches train th…