19 citations · 19 across the 2 of their papers we have counts for
3 papers
cs.AR2024
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
Hyungkyu Ham, Wonhyuk Yang, Yunseon Shin +5
As DNNs are widely adopted in various application domains while demanding increasingly higher compute and memory requirements, designing efficient and performant NPUs (Neural Proce…
cs.AR2024
Low-overhead General-purpose Near-Data Processing in CXL Memory Expanders
Hyungkyu Ham, Jeongmin Hong, Geonwoo Park +8
Emerging Compute Express Link (CXL) enables cost-efficient memory expansion beyond the local DRAM of processors. While its CXLmem protocol provides minimal latency overhead thro…
cs.AR2024★ 19 cited
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
Jeongmin Hong, Sungjun Cho, Geonwoo Park +3
We propose overcoming the memory capacity limitation of GPUs with high-capacity Storage-Class Memory (SCM) and DRAM cache. By significantly increasing the memory capacity with SCM,…