1 paper
Kunming Shao, Ming Zeng, Xin Yuan +5
Attention-FFN disaggregation maps LLM modules to specialized pools, creating an opening to keep Mixture-of-Experts (MoE) weights resident in a high-bandwidth FFN pool. Decode SLOs,…