1 paper
Kangyu Qiao, Shaolei Zhang, Yang Feng
With the growing computational demands of large language models (LLMs), efficient inference has become increasingly critical for practical deployment. Depth pruning has emerged as…