most citedSlice-and-Forge: Making Better Use of Caches for Graph Convolutional Network Accelerators

1 citations · 2 across the 9 of their papers we have counts for

collaborators

9 papers

cs.DC2024

PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices

Si Ung Noh, Junguk Hong, Chaemin Lim +5

Recent dual in-line memory modules (DIMMs) are starting to support processing-in-memory (PIM) by associating their memory banks with processing elements (PEs), allowing application…

cs.AR2024

Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System

Hongsun Jang, Jaeyong Song, Jaewon Jung +3

The recent huge advance of Large Language Models (LLMs) is mainly driven by the increase in the number of parameters. This has led to substantial memory capacity requirements, nece…

cs.DC2024

AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read Mapping

Seongyeon Park, Junguk Hong, Jaeyong Song +3

With the advance in genome sequencing technology, the lengths of deoxyribonucleic acid (DNA) sequencing results are rapidly increasing at lower prices than ever. However, the longe…

cs.LG2023

Pipe-BD: Pipelined Parallel Blockwise Distillation

Hongsun Jang, Jaewon Jung, Jaeyong Song +3

Training large deep neural network models is highly challenging due to their tremendous computational and memory requirements. Blockwise distillation provides one promising method…

cs.LG2023

SGCN: Exploiting Compressed-Sparse Features in Deep Graph Convolutional Network Accelerators

Mingi Yoo, Jaeyong Song, Jounghoo Lee +3

Graph convolutional networks (GCNs) are becoming increasingly popular as they overcome the limited applicability of prior neural networks. A GCN takes as input an arbitrarily struc…

cs.LG2023

Optimus-CC: Efficient Large NLP Model Training with 3D Parallelism Aware Communication Compression

Jaeyong Song, Jinkyu Yim, Jaewon Jung +4

In training of modern large natural language processing (NLP) models, it has become a common practice to split models using 3D parallelism to multiple GPUs. Such technique, however…