NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
arXiv:2403.00579 · doi:10.1145/3620666.3651380
Abstract
Modern transformer-based Large Language Models (LLMs) are constructed with a series of decoder blocks. Each block comprises three key components: (1) QKV generation, (2) multi-head attention, and (3) feed-forward networks. In batched processing, QKV generation and feed-forward networks involve compute-intensive matrix-matrix multiplications (GEMM), while multi-head attention requires bandwidth-heavy matrix-vector multiplications (GEMV). Machine learning accelerators like TPUs or NPUs are proficient in handling GEMM but are less efficient for GEMV computations. Conversely, Processing-in-Memory (PIM) technology is tailored for efficient GEMV computation, while it lacks the computational power to handle GEMM effectively. Inspired by this insight, we propose NeuPIMs, a heterogeneous acceleration system that jointly exploits a conventional GEMM-focused NPU and GEMV-optimized PIM devices. The main challenge in efficiently integrating NPU and PIM lies in enabling concurrent operations on both platforms, each addressing a specific kernel type. First, existing PIMs typically operate in a "blocked" mode, allowing only either NPU or PIM to be active at any given time. Second, the inherent dependencies between GEMM and GEMV in LLMs restrict their parallel processing. To tackle these challenges, NeuPIMs is equipped with dual row buffers in each bank, facilitating the simultaneous management of memory read/write operations and PIM commands. Further, NeuPIMs employs a runtime sub-batch interleaving technique to maximize concurrent execution, leveraging batch parallelism to allow two independent sub-batches to be pipelined within a single NeuPIMs device. Our evaluation demonstrates that compared to GPU-only, NPU-only, and a naïve NPU+PIM integrated acceleration approaches, NeuPIMs achieves 3, 2.4 and 1.6 throughput improvement, respectively.
16 pages, 15 figures
References in corpus (16)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- LLaMA: Open and Efficient Foundation Language Models
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Evaluating Large Language Models Trained on Code
- Training Compute-Optimal Large Language Models
- Code Llama: Open Foundation Models for Code
- TensorFlow-Serving: Flexible, High-Performance ML Serving
- Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep Learning
- GPT-NeoX-20B: An Open-Source Autoregressive Language Model
- Accelerating Recommendation System Training by Leveraging Popular Choices
- Pre-Trained Language Models for Interactive Decision-Making
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving
- TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings
- Accelerating Bandwidth-Bound Deep Learning Inference with Main-Memory Accelerators
- LightSeq: A High Performance Inference Library for Transformers
- Training Personalized Recommendation Systems from (GPU) Scratch: Look Forward not Backwards
Cited by in corpus (8)
- Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching
- Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
- How to keep pushing ML accelerator performance? Know your rooflines!
- Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
- Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
- Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
- AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization
- CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device