2 papers
cs.AR2025
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
Hyucksung Kwon, Kyungmo Koo, Janghyeon Kim +19
The expansion of long-context Large Language Models (LLMs) creates significant memory system challenges. While Processing-in-Memory (PIM) is a promising accelerator, we identify th…
cs.AR2024
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
Minseok Seo, Xuan Truong Nguyen, Seok Joong Hwang +18
Accelerating end-to-end inference of transformer-based large language models (LLMs) is a critical component of AI services in datacenters. However, diverse compute characteristics…