3 papers
cs.CV2025
Data-Free Group-Wise Fully Quantized Winograd Convolution via Learnable Scales
Shuokai Pan, Gerti Tuzi, Sudarshan Sreeram +1
Despite the revolutionary breakthroughs of large-scale text-to-image diffusion models for complex vision and downstream tasks, their extremely high computational and storage costs…
cs.LG2024
Highly Optimized Kernels and Fine-Grained Codebooks for LLM Inference on Arm CPUs
Dibakar Gope, David Mansell, Danny Loh +1
Large language models (LLMs) have transformed the way we think about language understanding and generation, enthralling both researchers and developers. However, deploying LLMs for…
cs.CV2024
Jumping through Local Minima: Quantization in the Loss Landscape of Vision Transformers
Natalia Frumkin, Dibakar Gope, Diana Marculescu
Quantization scale and bit-width are the most important parameters when considering how to quantize a neural network. Prior work focuses on optimizing quantization scales in a glob…