1 paper
Seungmin Oh, Donggeon Lee, Jongbin Ryu
Large language models achieve strong performance across diverse tasks, but deployment remains costly because of memory, latency, and energy demands. Structured pruning reduces thes…