1 paper
Minkyu Kim, Vincent-Daniel Yun, Youngrae Kim +3
Depth pruning improves the inference efficiency of large language models by removing Transformer blocks. Prior work typically treats layer redundancy as an inherent structural prop…