1 paper
Krishna Teja Chitty-Venkata, Sylvia Howland, Golara Azar +5
Mixture of Experts (MoE) models have enabled the scaling of Large Language Models (LLMs) and Vision Language Models (VLMs) by achieving massive parameter counts while maintaining c…