2 papers
cs.LG2026
Agentic Reinforcement Learning with Self-Distilled Reward Shaping
Ranxu Zhang, Guinan Chen, Chenshaodong +5
Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory-level rewards reveal success without identifying which intermediate decisions…
cs.LG2026
Latent Shadows: The Gaussian-Discrete Duality in Masked Diffusion
Guinan Chen, Xunpeng Huang, Ying Sun +3
Masked discrete diffusion is a dominant paradigm for high-quality language modeling where tokens are iteratively corrupted to a mask state, yet its inference efficiency is bottlene…