3 papers
cs.LG2026
AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization
Wanqi Yang, Yuexiao Ma, Alexander Conzelmann +4
Mixture-of-Experts (MoE) architectures scale model capacity through sparse expert activation, but their deployment remains memory-bound because all expert weights must reside in me…
cs.LG2026
Layer Collapse in Diffusion Language Models
Alexander Conzelmann, Albert Catalan-Tatjer, Shiwei Liu
Diffusion language models (DLMs) have recently emerged as competitive alternatives to autoregressive (AR) language models, yet differences in their activation dynamics remain poorl…
cs.LG2025
Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding
Alexander Conzelmann, Robert Bamler
The ever-growing size of neural networks poses serious challenges on resource-constrained devices, such as embedded sensors. Compression algorithms that reduce their size can mitig…