activity
20242026
collaborators

6 papers

cs.LG2026

FlowRL: A Taxonomy and Modular Framework for Reinforcement Learning with Diffusion Policies

Chenxiao Gao, Edward Chen, Tianyi Chen +1

Thanks to their remarkable flexibility, diffusion models and flow models have emerged as promising candidates for policy representation. However, efficient reinforcement learning (…

cs.LG2026

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning

Haitong Ma, Chenxiao Gao, Tianyi Chen +2

A commonly used family of RL algorithms for diffusion policies conducts softmax reweighting over samples from the behavior policy, which often induces an overgreedy policy and fail…

cs.LG2026

Revisiting DAgger in the Era of LLM-Agents

Changhao Li, Rushi Qiang, Jiawei Huang +4

Long-horizon LM agents learn from multi-turn interaction, where a single early mistake can alter the subsequent state distribution and derail the whole trajectory. Existing recipes…

cs.LG2026

Exploration-Driven Optimization for Test-Time Large Language Model Reasoning

Changhao Li, Yuchen Zhuang, Chenxiao Gao +4

Post-training techniques combined with inference-time scaling significantly enhance the reasoning and alignment capabilities of large language models (LLMs). However, a fundamental…

cs.LG2025

Reward Models in Deep Reinforcement Learning: A Survey

Rui Yu, Shenghua Wan, Yucen Wang +4

In reinforcement learning (RL), agents continually interact with the environment and use the feedback to refine their behavior. To guide policy optimization, reward models are intr…

cs.LG2024

Reinforced In-Context Black-Box Optimization

Lei Song, Chenxiao Gao, Ke Xue +5

Black-Box Optimization (BBO) has found successful applications in many fields of science and engineering. Recently, there has been a growing interest in meta-learning particular co…