SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
arXiv:2509.17072 · doi:10.1109/ASP-DAC66049.2026.11420607
Abstract
The growing scale of large language models (LLMs) has intensified demands on computation and memory, making efficient inference a key challenge. While sparsity can reduce these costs, existing design space exploration (DSE) frameworks often overlook compression formats, a key factor for leveraging sparsity on accelerators. This paper proposes SnipSnap, a joint compression format and dataflow co-optimization framework for efficient sparse LLM accelerator design. SnipSnap introduces: (1) a hierarchical compression format encoding to expand the design space; (2) an adaptive compression engine for selecting formats under diverse sparsity; and (3) a progressive co-search workflow that jointly optimizes dataflow and compression formats. SnipSnap achieves 18.24% average memory energy savings via format optimization, along with 2248.3 and 21.0 speedups over Sparseloop and DiMO-Sparse frameworks, respectively.
To appear in the 31st Asia and South Pacific Design Automation Conference (ASP-DAC 2026)
References in corpus (4)
- An Algorithm-Hardware Co-Optimized Framework for Accelerating N:M Sparse Transformers
- BitWave: Exploiting Column-Based Bit-Level Sparsity for Deep Learning Acceleration
- Efficient N:M Sparse DNN Training Using Algorithm, Architecture, and Dataflow Co-Design
- A 28nm 0.22μJ/token memory-compute-intensity-aware CNN-Transformer accelerator with hybrid-attention-based layer-fusion and cascaded pruning for semantic-segmentation