3 papers
cs.LG2026
Provable and Practical In-Context Policy Optimization for Self-Improvement
Tianrun Yu, Yuxiao Yang, Zhaoyang Wang +6
We study test-time scaling, where a model improves its answer through multi-round self-reflection at inference. We introduce In-Context Policy Optimization (ICPO), in which an agen…
cs.LG2025
Energy-Weighted Flow Matching for Offline Reinforcement Learning
Shiyuan Zhang, Weitong Zhang, Quanquan Gu
This paper investigates energy guidance in generative modeling, where the target distribution is defined as , with $…
cs.LG2024
Uncertainty-Aware Reward-Free Exploration with General Function Approximation
Junkai Zhang, Weitong Zhang, Dongruo Zhou +1
Mastering multiple tasks through exploration and learning in an environment poses a significant challenge in reinforcement learning (RL). Unsupervised RL has been introduced to add…