4 citations · 4 across the 4 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.LG2026
AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization
Wanqi Yang, Yuexiao Ma, Alexander Conzelmann +4
Mixture-of-Experts (MoE) architectures scale model capacity through sparse expert activation, but their deployment remains memory-bound because all expert weights must reside in me…
cs.LG2026
Layer Collapse in Diffusion Language Models
Alexander Conzelmann, Albert Catalan-Tatjer, Shiwei Liu
Diffusion language models (DLMs) have recently emerged as competitive alternatives to autoregressive (AR) language models, yet differences in their activation dynamics remain poorl…