1 paper · 1 filter
Aditya Vavre, Ethan He, Dennis Liu +4
Scaling large language models (LLMs) significantly improves performance but comes with prohibitive computational costs. Mixture-of-Experts (MoE) models offer an efficient alternati…