1 paper
Boxiang Zhang, Baijian Yang
Transformers achieve strong accuracy but incur high compute and memory cost. Structured pruning reduces inference cost, but most methods rely on retraining or multi-stage optimizat…