OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization
arXiv:2304.07493 · doi:10.1145/3579371.3589038
Abstract
Transformer-based large language models (LLMs) have achieved great success with the growing model size. LLMs' size grows by every two years, which outpaces the hardware progress and makes model inference increasingly costly. Model quantization is a promising approach to mitigate the widening gap between LLM size and hardware capacity. However, the existence of outliers, values with significant magnitudes, in LLMs makes existing quantization methods less effective. Prior outlier-aware quantization schemes adopt sparsity encoding techniques to separate outliers from normal values where the process requires global coordination (e.g., a global sparsity coordination list). This incurs complex encoding/decoding hardware logics and an extra orchestration controller for the computation between outlier and normal values. As such, it is not hardware-efficient and hence only achieves sub-optimal quantization benefits. We propose OliVe, an algorithm/architecture co-designed solution that adopts an outlier-victim pair (OVP) quantization and handles outlier values locally with low hardware overheads and high performance gains. The key insight of OliVe is that outliers are important while the normal values next to them are not. Thus those normal values (called victims) can be sacrificed to accommodate outliers. This enables a memory-aligned OVP encoding scheme, which can be efficiently integrated to the existing hardware accelerators like systolic array and tensor core. As a result, OliVe-based accelerator surpasses the existing outlier-aware accelerator, GOBO, by 4.5 speedup and 4.0 energy reduction, respectively, with a superior model accuracy.
ISCA 2023
References in corpus (11)
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- Deep Learning with Limited Numerical Precision
- PACT: Parameterized Clipping Activation for Quantized Neural Networks
- Exploring Modern GPU Memory System Design Challenges through Accurate Modeling
- DeepScaleTool : A Tool for the Accurate Estimation of Technology Scaling in the Deep-Submicron Era
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- Outlier Suppression: Pushing the Limit of Low-bit Transformer Language Models
- SQuant: On-the-Fly Data-Free Quantization via Diagonal Hessian Approximation
- Characterizing and Demystifying the Implicit Convolution Algorithm on Commercial Matrix-Multiplication Accelerators
- Efficient Adaptive Activation Rounding for Post-Training Quantization
Cited by in corpus (11)
- Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges
- Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching
- Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
- LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
- Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
- VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
- FuseMax: Leveraging Extended Einsums to Optimize Attention Accelerator Design
- Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
- FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor Management
- GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference
- LightNobel: Improving Sequence Length Limitation in Protein Structure Prediction Model via Adaptive Activation Quantization