2 papers
cs.AR2024
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
Minseok Seo, Xuan Truong Nguyen, Seok Joong Hwang +18
Accelerating end-to-end inference of transformer-based large language models (LLMs) is a critical component of AI services in datacenters. However, diverse compute characteristics…
cs.CL2024
Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization
Janghwan Lee, Minsoo Kim, Seungcheol Baek +3
Large Language Models (LLMs) are proficient in natural language processing tasks, but their deployment is often restricted by extensive parameter sizes and computational demands. T…