Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Reasoning with Latent Tokens in Diffusion Language Models
Andre He, Sean Welleck, Daniel Fried
Discrete diffusion models have recently become competitive with autoregressive models for language modeling, even outperforming them on reasoning tasks requiring planning and globa…
cs.LG2025
Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening
Andre He, Daniel Fried, Sean Welleck
Reinforcement learning is emerging as a primary driver for improving language model reasoning capabilities. A fundamental question is whether current reinforcement learning algorit…