activity
20202026
most citedFrequency-based Search-control in Dyna

2 citations · 4 across the 9 of their papers we have counts for

collaborators

13 papers

cs.LG2026

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic

Jiamin He, Samuel Neumann, Jincheng Mei +2

Mixture policies theoretically offer greater flexibility than unimodal policies in continuous action reinforcement learning, but the practical benefits of this complexity remain el…

cs.LG2025

Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation

Max Qiushi Lin, Jincheng Mei, Matin Aghaei +6

Policy gradient (PG) methods have played an essential role in the empirical successes of reinforcement learning. In order to handle large state-action spaces, PG methods are typica…

cs.LG2025

Ordering-based Conditions for Global Convergence of Policy Gradient Methods

Jincheng Mei, Bo Dai, Alekh Agarwal +3

We prove that, for finite-arm bandits with linear function approximation, the global convergence of policy gradient (PG) methods depends on inter-related properties between the pol…

cs.LG2025

Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates

Jincheng Mei, Bo Dai, Alekh Agarwal +4

We provide a new understanding of the stochastic gradient bandit algorithm by showing that it converges to a globally optimal policy almost surely using \emph{any} constant learnin…

cs.LG2024

Faster WIND: Accelerating Iterative Best-of- Distillation for LLM Alignment

Tong Yang, Jincheng Mei, Hanjun Dai +5

Recent advances in aligning large language models with human preferences have corroborated the growing importance of best-of-N distillation (BOND). However, the iterative BOND algo…

cs.LG2024

Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function Approximation

Fengdi Che, Chenjun Xiao, Jincheng Mei +6

We prove that the combination of a target network and over-parameterized linear function approximation establishes a weaker convergence condition for bootstrapped value estimation…