32 citations · 77 across the 36 of their papers we have counts for
28 papers · 1 filter
Temporal Self-Distillation: Faster Inference in Discrete Diffusion Language Models
Shijian Xu, Andrea Miele, Metod Jazbec +3
Diffusion language models (dLLMs) promise fast inference by generating multiple tokens in parallel, but suffer severe performance degradation when parallel decoding is pushed too a…
Re-evaluating Confidence Remasking in Masked Diffusion Language Models
Stipe Frkovic, Metod Jazbec, Dan Zhang +3
Masked diffusion language models (dLLMs) have recently emerged as a competitive alternative to autoregressive language models, with the promise of faster inference via parallel tok…
PROWL: Prioritized Regret-Driven Optimization for World Model Learning
Ahmet H. Güzel, Jenny Seidenschwarz, Benjamin Graham +3
Modern action-conditioned video world models achieve strong short-horizon visual realism, yet remain unreliable on rare, interaction-critical transitions that dominate downstream p…
GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models
Xiaohang Tang, Keyue Jiang, Che Liu +4
Reinforcement learning (RL) can be used to improve the policy (denoiser) of diffusion large language models (dLLMs), while being hindered by the intractability of the policy likeli…
wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
Xiaohang Tang, Rares Dolga, Sangwoong Yoon +1
Improving the reasoning capabilities of diffusion-based large language models (dLLMs) through reinforcement learning (RL) remains an open problem. The intractability of dLLMs likel…
Robust Multi-Objective Controlled Decoding of Large Language Models
Seongho Son, William Bankes, Sangwoong Yoon +3
We introduce Robust Multi-Objective Decoding (RMOD), a novel inference-time algorithm that robustly aligns Large Language Models (LLMs) to multiple human objectives (e.g., instruct…