3 papers
cs.AR2026
Exploring High-Bandwidth Flash for Modern LLM Inference: Opportunities and Challenges
Dowon Son, Yonggon Park, Hyunuk Cho +5
This work investigates the potential benefits and technical challenges of using high-bandwidth flash (HBF) for large language model (LLM) inference. HBF has gained increasing atten…
cs.AR2026
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention
Hyungkyu Ham, Junhyeong Bae, Seungheon Lee +2
This paper presents a heterogeneous decode-phase serving system that relocates the KV cache out of GPU memory, motivated by the retrieval-based sparse attention that recent frontie…
cs.AR2024
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
Hyungkyu Ham, Wonhyuk Yang, Yunseon Shin +5
As DNNs are widely adopted in various application domains while demanding increasingly higher compute and memory requirements, designing efficient and performant NPUs (Neural Proce…