2 papers
cs.AR2025
Dissecting and Re-architecting 3D NAND Flash PIM Arrays for Efficient Single-Batch Token Generation in LLMs
Yongjoo Jang, Sangwoo Hwang, Hojin Lee +4
The advancement of large language models has led to models with billions of parameters, significantly increasing memory and compute demands. Serving such models on conventional har…
cs.AR2025
All-rounder: A Flexible AI Accelerator with Diverse Data Format Support and Morphable Structure for Multi-DNN Processing
Seock-Hwan Noh, Seungpyo Lee, Banseok Shin +3
Recognizing the explosive increase in the use of AI-based applications, several industrial companies developed custom ASICs (e.g., Google TPU, IBM RaPiD, Intel NNP-I/NNP-T) and con…