Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
arXiv:2312.15234 · doi:10.1145/3754448
Abstract
In the rapidly evolving landscape of artificial intelligence (AI), generative large language models (LLMs) stand at the forefront, revolutionizing how we interact with our data. However, the computational intensity and memory consumption of deploying these models present substantial challenges in terms of serving efficiency, particularly in scenarios demanding low latency and high throughput. This survey addresses the imperative need for efficient LLM serving methodologies from a machine learning system (MLSys) research perspective, standing at the crux of advanced AI innovations and practical system optimizations. We provide in-depth analysis, covering a spectrum of solutions, ranging from cutting-edge algorithmic modifications to groundbreaking changes in system designs. The survey aims to provide a comprehensive understanding of the current state and future directions in efficient LLM serving, offering valuable insights for researchers and practitioners in overcoming the barriers of effective LLM deployment, thereby reshaping the future of AI.
ACM Computing Surveys
References in corpus (19)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- MLP-Mixer: An all-MLP Architecture for Vision
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers
- Mixture-of-Experts with Expert Choice Routing
- Efficiently Scaling Transformer Inference
- Reducing Activation Recomputation in Large Transformer Models
- Hash Layers For Large Sparse Models
- Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques and Tools
- FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement
- Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding
- Tutel: Adaptive Mixture-of-Experts at Scale
- Accelerating Transformer Inference for Translation via Parallel Decoding
- Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
- STI: Turbocharge NLP Inference at the Edge via Elastic Pipelining
- MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
- Relax: Composable Abstractions for End-to-End Dynamic Machine Learning