1 paper
Qingyuan Liu, Liyan Chen, Yanning Yang +6
Attention-FC Disaggregated (AFD) LLM inference systems offload memory-bound Attention operations to memory-rich accelerators (e.g., CPUs, HBM-PIM) while retaining compute-bound Ful…