2 papers
cs.LG2026
Search Your Block Floating Point Scales!
Tanmaey Gupta, Hayden Prairie, Xiaoxia Wu +10
Quantization has emerged as a standard technique for accelerating inference for generative models by enabling faster low-precision computations and reduced memory transfers. Recent…
cs.LG2026
DUEL: Exact Likelihood for Masked Diffusion via Deterministic Unmasking
Gilad Turok, Chris De Sa, Volodymyr Kuleshov
Masked diffusion models (MDMs) generate text by iteratively selecting positions to unmask and then predicting tokens at those positions. Yet MDMs lack proper likelihood evaluation:…