6 papers
FlowRL: A Taxonomy and Modular Framework for Reinforcement Learning with Diffusion Policies
Chenxiao Gao, Edward Chen, Tianyi Chen +1
Thanks to their remarkable flexibility, diffusion models and flow models have emerged as promising candidates for policy representation. However, efficient reinforcement learning (…
GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning
Haitong Ma, Chenxiao Gao, Tianyi Chen +2
A commonly used family of RL algorithms for diffusion policies conducts softmax reweighting over samples from the behavior policy, which often induces an overgreedy policy and fail…
Revisiting DAgger in the Era of LLM-Agents
Changhao Li, Rushi Qiang, Jiawei Huang +4
Long-horizon LM agents learn from multi-turn interaction, where a single early mistake can alter the subsequent state distribution and derail the whole trajectory. Existing recipes…
Exploration-Driven Optimization for Test-Time Large Language Model Reasoning
Changhao Li, Yuchen Zhuang, Chenxiao Gao +4
Post-training techniques combined with inference-time scaling significantly enhance the reasoning and alignment capabilities of large language models (LLMs). However, a fundamental…
Reward Models in Deep Reinforcement Learning: A Survey
Rui Yu, Shenghua Wan, Yucen Wang +4
In reinforcement learning (RL), agents continually interact with the environment and use the feedback to refine their behavior. To guide policy optimization, reward models are intr…
Reinforced In-Context Black-Box Optimization
Lei Song, Chenxiao Gao, Ke Xue +5
Black-Box Optimization (BBO) has found successful applications in many fields of science and engineering. Recently, there has been a growing interest in meta-learning particular co…