1 paper
Kai Liu, Bowen Xu, Shaoyu Wu +4
Activation sparsity can reduce the computational overhead and memory transfers during the forward pass of Large Language Model (LLM) inference. Existing methods face limitations, e…