1 paper
Weiyu Zhou, Chen Ding, Mingyuan Liu +7
Large language model (LLM) inference suffers from growing inefficiency across the prefill and decode phases, especially under weight-only quantization, where activations remain in…