1 paper
Udbhav Bamba, Arnav Chavan, Aryamaan Thakur +2
The scaling of Large Language Models (LLMs) has driven significant performance gains but created substantial challenges in inference efficiency. While Mixture of Experts (MoEs) arc…