2 papers
cs.AR2026
MX-SAFE: Versatile Inference- and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation
Dahoon Park, Jahyun Koo, Sangwoo Hwang +1
As the demand for deep learning grows, cost reduction through quantization has become essential for both training and inference. In 2022, the Open Compute Project (OCP) consortium…
cs.AR2025
Dissecting and Re-architecting 3D NAND Flash PIM Arrays for Efficient Single-Batch Token Generation in LLMs
Yongjoo Jang, Sangwoo Hwang, Hojin Lee +4
The advancement of large language models has led to models with billions of parameters, significantly increasing memory and compute demands. Serving such models on conventional har…