3 papers
cs.LG2026
Displacement Is Not Direction: Evaluating Fidelity Metrics for Quantized LLM Deployment
MiloÅ¡ NikoliÄ, Ali Hadi Zadeh, Enrique Torres Sanchez +1
Fidelity metrics, such as per-token KL divergence (KLD) against a high-precision reference, are often used in practice as low-cost proxies for benchmark quality. We test this pract…
cs.CV2024
Low-Bitwidth Floating Point Quantization for Efficient High-Quality Diffusion Models
Cheng Chen, Christina Giannoula, Andreas Moshovos
Diffusion models are emerging models that generate images by iteratively denoising random Gaussian noise using deep neural networks. These models typically exhibit high computation…
cs.LG2024
Schrödinger's FP: Dynamic Adaptation of Floating-Point Containers for Deep Learning Training
MiloÅ¡ NikoliÄ, Enrique Torres Sanchez, Jiahui Wang +5
The transfer of tensors from/to memory during neural network training dominates time and energy. To improve energy efficiency and performance, research has been exploring ways to u…