Relax: Composable Abstractions for End-to-End Dynamic Machine Learning
arXiv:2311.02103 · doi:10.1145/3676641.3716249
Abstract
Dynamic shape computations have become critical in modern machine learning workloads, especially in emerging large language models. The success of these models has driven the demand for their universal deployment across a diverse set of backend environments. In this paper, we present Relax, a compiler abstraction for optimizing end-to-end dynamic machine learning workloads. Relax introduces a cross-level abstraction that encapsulates computational graphs, loop-level tensor programs, and external library calls in a single representation. Relax also introduces first-class symbolic shape annotations to track dynamic shape computations globally across the program, enabling dynamic shape-aware cross-level optimizations. We build an end-to-end compilation framework using the proposed approach to optimize dynamic shape models. Experimental results on LLMs show that Relax delivers performance competitive with state-of-the-art systems across various GPUs and enables deployment of emerging models to a broader set of emerging environments, including mobile phones, embedded devices, and web browsers.
To appear at ASPLOS 2025 (16 pages, 20 figures)
References in corpus (14)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Learning Transferable Visual Models From Natural Language Supervision
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Robust Speech Recognition via Large-Scale Weak Supervision
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- Code Llama: Open Foundation Models for Code
- GPT-NeoX-20B: An Open-Source Autoregressive Language Model
- Torch.fx: Practical Program Capture and Transformation for Deep Learning in Python
- WhisperX: Time-Accurate Speech Transcription of Long-Form Audio
- Tensor Program Optimization with Probabilistic Programs
- Stream-K: Work-centric Parallel Decomposition for Dense Matrix-Matrix Multiplication on the GPU
- The CoRa Tensor Compiler: Compilation for Ragged Tensors with Minimal Padding
- DISC: A Dynamic Shape Compiler for Machine Learning Workloads
- Axon: A Language for Dynamic Shapes in Deep Learning Graphs