1 paper
Maosen Zhao, Pengtao Chen, Chong Yu +3
Model quantization reduces the bit-width of weights and activations, improving memory efficiency and inference speed in diffusion models. However, achieving 4-bit quantization rema…