1 citations · 1 across the 4 of their papers we have counts for
3 papers · 1 filter
Shieldstral
Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli +274
We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7 its size on text safety benchmarks and set…
Ministral 3
Alexander H. Liu, Kartik Khandelwal, Sandeep Subramanian +116
We introduce the Ministral 3 series, a family of parameter-efficient dense language models designed for compute and memory constrained applications, available in three model sizes:…
Mixture of Tokens: Continuous MoE through Cross-Example Aggregation
Szymon Antoniak, MichaŠKrutul, Maciej Pióro +7
Mixture of Experts (MoE) models based on Transformer architecture are pushing the boundaries of language and vision tasks. The allure of these models lies in their ability to subst…