5 citations · 9 across the 11 of their papers we have counts for
7 papers · 1 filter
Pipe-BD: Pipelined Parallel Blockwise Distillation
Hongsun Jang, Jaewon Jung, Jaeyong Song +3
Training large deep neural network models is highly challenging due to their tremendous computational and memory requirements. Blockwise distillation provides one promising method…
SGCN: Exploiting Compressed-Sparse Features in Deep Graph Convolutional Network Accelerators
Mingi Yoo, Jaeyong Song, Jounghoo Lee +3
Graph convolutional networks (GCNs) are becoming increasingly popular as they overcome the limited applicability of prior neural networks. A GCN takes as input an arbitrarily struc…
Optimus-CC: Efficient Large NLP Model Training with 3D Parallelism Aware Communication Compression
Jaeyong Song, Jinkyu Yim, Jaewon Jung +4
In training of modern large natural language processing (NLP) models, it has become a common practice to split models using 3D parallelism to multiple GPUs. Such technique, however…
Slice-and-Forge: Making Better Use of Caches for Graph Convolutional Network Accelerators
Mingi Yoo, Jaeyong Song, Hyeyoon Lee +4
Graph convolutional networks (GCNs) are becoming increasingly popular as they can process a wide variety of data formats that prior deep neural networks cannot easily support. One…
Enabling Hard Constraints in Differentiable Neural Network and Accelerator Co-Exploration
Deokki Hong, Kanghyun Choi, Hye Yoon Lee +4
Co-exploration of an optimal neural architecture and its hardware accelerator is an approach of rising interest which addresses the computational cost problem, especially in low-pr…
Qimera: Data-free Quantization with Synthetic Boundary Supporting Samples
Kanghyun Choi, Deokki Hong, Noseong Park +2
Model quantization is known as a promising method to compress deep neural networks, especially for inferences on lightweight mobile or edge devices. However, model quantization usu…