6 citations · 8 across the 3 of their papers we have counts for
1 paper · 1 filter
Yuseon Choi, Jingu Lee, Jungjun Oh +5
Mixture-of-Experts (MoE) models have become the dominant architecture for large-scale language models, yet on-premises serving remains fundamentally memory-bound as batching turns…