10 citations · 22 across the 24 of their papers we have counts for
3 papers · 2 filters
When to Prune? A Policy towards Early Structural Pruning
Maying Shen, Pavlo Molchanov, Hongxu Yin +1
Pruning enables appealing reductions in network memory footprint and time complexity. Conventional post-training pruning techniques lean towards efficient inference while overlooki…
HALP: Hardware-Aware Latency Pruning
Maying Shen, Hongxu Yin, Pavlo Molchanov +3
Structural pruning can simplify network architecture and improve inference speed. We propose Hardware-Aware Latency Pruning (HALP) that formulates structural pruning as a global re…
Global Vision Transformer Pruning with Hessian-Aware Saliency
Huanrui Yang, Hongxu Yin, Maying Shen +3
Transformers yield state-of-the-art results across many tasks. However, their heuristically designed architecture impose huge computational costs during inference. This work aims o…