1 paper
Yuhua Zhou, Shaoqi Yu, Shichao Weng +4
Large language models (LLMs) incur high inference cost due to their depth and parameter scale. Depth pruning can reduce latency by skipping redundant Transformer blocks, but existi…