4 citations · 4 across the 3 of their papers we have counts for
2 papers
cs.LG2026
AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization
Wanqi Yang, Yuexiao Ma, Alexander Conzelmann +4
Mixture-of-Experts (MoE) architectures scale model capacity through sparse expert activation, but their deployment remains memory-bound because all expert weights must reside in me…
cs.LG2025
Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding
Alexander Conzelmann, Robert Bamler
The ever-growing size of neural networks poses serious challenges on resource-constrained devices, such as embedded sensors. Compression algorithms that reduce their size can mitig…