2 papers
cs.AR2026
FlashAccel: Leveraging High-Bandwidth Flash (HBF) for High-Throughput LLM Inference
Xinyu Wang, Yalong Xue, Xiaotian Sun +5
Large language model (LLM) inference is increasingly limited by the capacity of High-Bandwidth Memory (HBM) in GPUs, as model weights and KV cache grow rapidly. High-Bandwidth Flas…
cs.AR2024
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
Xiaotian Sun, Xinyu Wang, Wanqian Li +2
Various processing-in-memory (PIM) accelerators based on various devices, micro-architectures, and interfaces have been proposed to accelerate deep neural networks (DNNs). How to d…