266 citations · 269 across the 8 of their papers we have counts for
8 papers
Pipe-BD: Pipelined Parallel Blockwise Distillation
Hongsun Jang, Jaewon Jung, Jaeyong Song +3
Training large deep neural network models is highly challenging due to their tremendous computational and memory requirements. Blockwise distillation provides one promising method…
SGCN: Exploiting Compressed-Sparse Features in Deep Graph Convolutional Network Accelerators
Mingi Yoo, Jaeyong Song, Jounghoo Lee +3
Graph convolutional networks (GCNs) are becoming increasingly popular as they overcome the limited applicability of prior neural networks. A GCN takes as input an arbitrarily struc…
Optimus-CC: Efficient Large NLP Model Training with 3D Parallelism Aware Communication Compression
Jaeyong Song, Jinkyu Yim, Jaewon Jung +4
In training of modern large natural language processing (NLP) models, it has become a common practice to split models using 3D parallelism to multiple GPUs. Such technique, however…
Slice-and-Forge: Making Better Use of Caches for Graph Convolutional Network Accelerators
Mingi Yoo, Jaeyong Song, Hyeyoon Lee +4
Graph convolutional networks (GCNs) are becoming increasingly popular as they can process a wide variety of data formats that prior deep neural networks cannot easily support. One…
Enabling Hard Constraints in Differentiable Neural Network and Accelerator Co-Exploration
Deokki Hong, Kanghyun Choi, Hye Yoon Lee +4
Co-exploration of an optimal neural architecture and its hardware accelerator is an approach of rising interest which addresses the computational cost problem, especially in low-pr…
SaLoBa: Maximizing Data Locality and Workload Balance for Fast Sequence Alignment on GPUs
Seongyeon Park, Hajin Kim, Tanveer Ahmad +5
Sequence alignment forms an important backbone in many sequencing applications. A commonly used strategy for sequence alignment is an approximate string matching with a two-dimensi…