27 citations · 137 across the 49 of their papers we have counts for
3 papers · 2 filters
ReGuide: From Test-Time Guidance to Self-Improving Diffusion Policies
Tzu-Hsiang Lin, Srinivas Shakkottai, Dileep Kalathil +1
Behavior-cloned diffusion policies are expressive but remain vulnerable to covariate shift: small deviations from demonstrated states can compound into task failure. Existing metho…
Reinforcement Learning for Diffusion LLMs with Entropy-Guided Step Selection and Stepwise Advantages
Vishnu Teja Kunde, Fatemeh Doudi, Mahdi Farahbakhsh +3
Reinforcement learning (RL) has been effective for post-training autoregressive (AR) language models, but extending these methods to diffusion language models (DLMs) is challenging…
Optimistic World Models: Efficient Exploration in Model-Based Deep Reinforcement Learning
Akshay Mete, Shahid Aamir Sheikh, Tzu-Hsiang Lin +2
Efficient exploration remains a central challenge in reinforcement learning (RL), particularly in sparse-reward environments. We introduce Optimistic World Models (OWMs), a princip…