10 papers
DSL-LLaDA: Scaling Continuous Denoising to 8B Masked Diffusion LMs
Longxuan Yu, Yunshu Wu, Yu Fu +5
Discrete Masked diffusion language models generate text by iterative parallel decoding, but few-step decoding suffers from a tradeoff between length and quality: with a fixed step…
Local MAP Sampling for Diffusion Models
Shaorong Zhang, Rob Brekelmans, Greg Ver Steeg
Diffusion Posterior Sampling (DPS) provides a principled Bayesian approach to inverse problems by sampling from . While posterior sampling is valuable for capturing…
Discrete Stochastic Localization for Non-autoregressive Generation
Yunshu Wu, Jiayi Cheng, Longxuan Yu +4
Continuous diffusion is a natural framework for non-autoregressive generation but has generally lagged behind masked discrete diffusion models (MDMs) on discrete sequence generatio…
Discrete Stochastic Localization for Non-autoregressive Generation
Yunshu Wu, Jiayi Cheng, Longxuan Yu +4
Continuous diffusion is a natural framework for non-autoregressive generation but has generally lagged behind masked discrete diffusion models (MDMs) on discrete sequence generatio…
Scalable Spatio-Temporal SE(3) Diffusion for Long-Horizon Protein Dynamics
Nima Shoghi, Yuxuan Liu, Yuning Shen +3
Molecular dynamics (MD) simulations remain the gold standard for studying protein dynamics, but their computational cost limits access to biologically relevant timescales. Recent g…
Generation Order and Parallel Decoding in Masked Diffusion Models: An Information-Theoretic Perspective
Shaorong Zhang, Longxuan Yu, Rob Brekelmans +3
Masked Diffusion Models (MDMs) significantly accelerate inference by trading off sequential determinism. However, the theoretical mechanisms governing generation order and the risk…