3 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.LG2024★ 3 cited
Scaling Laws for Fine-Grained Mixture of Experts
Jakub Krajewski, Jan Ludziejewski, Kamil Adamczewski +9
Mixture of Experts (MoE) models have emerged as a primary solution for reducing the computational cost of Large Language Models. In this work, we analyze their scaling properties,…
cs.CL2023★ 1 cited
Mixture of Tokens: Continuous MoE through Cross-Example Aggregation
Szymon Antoniak, Michał Krutul, Maciej Pióro +7
Mixture of Experts (MoE) models based on Transformer architecture are pushing the boundaries of language and vision tasks. The allure of these models lies in their ability to subst…