1 paper
Pratyush Dhingra, Pramit Kumar Pal, Janardhan Rao Doppa +1
Mixture of Experts (MoE) architectures have emerged as a dominant paradigm for scaling Large Language Models (LLMs). However, MoE inference on conventional hardware is constrained…