2 papers
cs.LG2025
Quant-Trim in Practice: Improved Cross-Platform Low-Bit Deployment on Edge NPUs
Rayen Dhahri, Steffen Urban
Specialized edge accelerators rely on low-bit quantization, but vendor compilers differ in scaling, clipping, and kernel support, often as black boxes. The same floating-point (FP)…
cs.LG2024
Shaving Weights with Occam's Razor: Bayesian Sparsification for Neural Networks Using the Marginal Likelihood
Rayen Dhahri, Alexander Immer, Betrand Charpentier +2
Neural network sparsification is a promising avenue to save computational time and memory costs, especially in an age where many successful AI models are becoming too large to naï…