1 paper
Amirmohsen Sattarifard, Sepehr Lavasani, Kunlin Zhang +5
Inference-time sparsification is a promising path to deploy large language models (LLMs) on resource-constrained devices, yet existing training-free methods typically estimate feed…