2 papers
cs.LG2026
Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy
Haozhe Hu, Hao Wu, Anhao Zhao +4
Pruning has emerged as a dominant paradigm for accelerating large language model (LLM) inference, spanning a broad spectrum of methods that remove computation across tokens, layers…
cs.LG2026
From LLMs to LRMs: Rethinking Pruning for Reasoning-Centric Models
Longwei Ding, Anhao Zhao, Fanghua Ye +2
Large language models (LLMs) are increasingly costly to deploy, motivating extensive research on model pruning. However, most existing studies focus on instruction-following LLMs,…