5 citations · 5 across the 4 of their papers we have counts for
3 papers · 1 filter
Pipe-BD: Pipelined Parallel Blockwise Distillation
Hongsun Jang, Jaewon Jung, Jaeyong Song +3
Training large deep neural network models is highly challenging due to their tremendous computational and memory requirements. Blockwise distillation provides one promising method…
Enabling Hard Constraints in Differentiable Neural Network and Accelerator Co-Exploration
Deokki Hong, Kanghyun Choi, Hye Yoon Lee +4
Co-exploration of an optimal neural architecture and its hardware accelerator is an approach of rising interest which addresses the computational cost problem, especially in low-pr…
NN-LUT: Neural Approximation of Non-Linear Operations for Efficient Transformer Inference
Joonsang Yu, Junki Park, Seongmin Park +4
Non-linear operations such as GELU, Layer normalization, and Softmax are essential yet costly building blocks of Transformer models. Several prior works simplified these operations…