2 papers
cs.LG2024
IM-Unpack: Training and Inference with Arbitrarily Low Precision Integers
Zhanpeng Zeng, Karthikeyan Sankaralingam, Vikas Singh
GEneral Matrix Multiply (GEMM) is a central operation in deep learning and corresponds to the largest chunk of the compute footprint. Therefore, improving its efficiency is an acti…
cs.LG2024
LookupFFN: Making Transformers Compute-lite for CPU inference
Zhanpeng Zeng, Michael Davies, Pranav Pulijala +2
While GPU clusters are the de facto choice for training large deep neural network (DNN) models today, several reasons including ease of workflow, security and cost have led to effo…