1 paper
Hanzhen Wang, Jiaming Xu, Yushun Xiang +4
Pruning is a typical acceleration technique for compute-bound models by removing computation on unimportant values. Recently, it has been applied to accelerate Vision-Language-Acti…