5 papers
What Does a Discrete Diffusion Model Learn?
Rodrigo Casado Noguerales, Bernhard Schölkopf, Thomas Hofmann +1
What does a discrete diffusion model learn: a denoiser, a score ratio, or a bridge plug-in predictor? At the level of jump rates, these are one object in different coordinates, and…
Intrinsically Interpretable Attention via Sparse Post-Training
Florent Draye, Anson Lei, Hsiao-Ru Pan +2
We introduce a simple post-training method that makes transformer attention sparse without sacrificing performance. Applying a flexible sparsity regularisation under a constrained-…
Scaling Behavior of Discrete Diffusion Language Models
Dimitri von Rütte, Janis Fluri, Omead Pooladzandi +3
Modern LLM pre-training consumes vast amounts of compute and training data, making the scaling behavior, or scaling laws, of different models a key distinguishing factor. Discrete…
DiffRatio: Training One-Step Diffusion Models Without Teacher Supervision
Wenlin Chen, Mingtian Zhang, Jiajun He +4
Score-based distillation methods (e.g., variational score distillation) train one-step diffusion models by first pre-training a teacher score model and then distilling it into a on…
Generalized Interpolating Discrete Diffusion
Dimitri von Rütte, Janis Fluri, Yuhui Ding +3
While state-of-the-art language models achieve impressive results through next-token prediction, they have inherent limitations such as the inability to revise already generated to…