4 papers
AutoGNN: End-to-End Hardware-Driven Graph Preprocessing for Enhanced GNN Performance
Seungkwan Kang, Seungjun Lee, Donghyun Gouk +9
Graph neural network (GNN) inference faces significant bottlenecks in preprocessing, which often dominate overall inference latency. We introduce AutoGNN, an FPGA-based accelerator…
MPI-over-CXL: Enhancing Communication Efficiency in Distributed HPC Systems
Miryeong Kwon, Donghyun Gouk, Hyein Woo +10
MPI implementations commonly rely on explicit memory-copy operations, incurring overhead from redundant data movement and buffer management. This overhead notably impacts HPC workl…
From Block to Byte: Transforming PCIe SSDs with CXL Memory Protocol and Instruction Annotation
Miryeong Kwon, Donghyun Gouk, Junhyeok Jang +8
This paper explores how Compute Express Link (CXL) can transform PCIe-based block storage into a scalable, byte-addressable working memory. We address the challenges of adapting bl…
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
Donghyun Gouk, Seungkwan Kang, Seungjun Lee +8
This work introduces a GPU storage expansion solution utilizing CXL, featuring a novel GPU system design with multiple CXL root ports for integrating diverse storage media (DRAMs a…