3 papers
cs.LG2025
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
Ahmed F. AbouElhamayed, Jordan Dotzel, Yash Akhauri +6
Large language models have high compute, latency, and memory requirements. While specialized accelerators such as GPUs and TPUs typically run these workloads, CPUs are more widely…
cs.LG2024
Post-Training Statistical Calibration for Higher Activation Sparsity
Vui Seng Chua, Yujie Pan, Nilesh Jain
We present Statistical Calibrated Activation Pruning (SCAP), a post-training activation pruning framework that (1) generalizes sparsification by input activations of Fully-Connecte…
cs.NE2021
Neuroevolution-Enhanced Multi-Objective Optimization for Mixed-Precision Quantization
Santiago Miret, Vui Seng Chua, Mattias Marder +3
Mixed-precision quantization is a powerful tool to enable memory and compute savings of neural network workloads by deploying different sets of bit-width precisions on separate com…