2 papers
cs.LG2024
OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
Jahyun Koo, Dahoon Park, Sangwoo Jung +1
To overcome the burden on the memory size and bandwidth due to ever-increasing size of large language models (LLMs), aggressive weight quantization has been recently studied, while…
cs.AR2022
LightNorm: Area and Energy-Efficient Batch Normalization Hardware for On-Device DNN Training
Seock-Hwan Noh, Junsang Park, Dahoon Park +3
When training early-stage deep neural networks (DNNs), generating intermediate features via convolution or linear layers occupied most of the execution time. Accordingly, extensive…