3 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.LG2025
On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures
Ming Yin, Mengdi Wang, Yu-Xiang Wang
This article reviews the recent advances on the statistical foundation of reinforcement learning (RL) in the offline and low-adaptive settings. We will start by arguing why offline…
cs.CL2024
Fast Best-of-N Decoding via Speculative Rejection
Hanshi Sun, Momin Haider, Ruiqi Zhang +6
The safe and effective deployment of Large Language Models (LLMs) involves a critical step called alignment, which ensures that the model's responses are in accordance with human p…
cs.LG2022★ 3 cited
Near-optimal Offline Reinforcement Learning with Linear Representation: Leveraging Variance Information with Pessimism
Ming Yin, Yaqi Duan, Mengdi Wang +1
Offline reinforcement learning, which seeks to utilize offline/historical data to optimize sequential decision-making strategies, has gained surging prominence in recent studies. D…