2 papers
cs.AR2025
Dissecting and Re-architecting 3D NAND Flash PIM Arrays for Efficient Single-Batch Token Generation in LLMs
Yongjoo Jang, Sangwoo Hwang, Hojin Lee +4
The advancement of large language models has led to models with billions of parameters, significantly increasing memory and compute demands. Serving such models on conventional har…
cs.LG2024
OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
Jahyun Koo, Dahoon Park, Sangwoo Jung +1
To overcome the burden on the memory size and bandwidth due to ever-increasing size of large language models (LLMs), aggressive weight quantization has been recently studied, while…