7 papers
On the Complexity of Offline Reinforcement Learning with -Approximation and Partial Coverage
Haolin Liu, Braham Snyder, Chen-Yu Wei
We study offline reinforcement learning under -approximation and partial coverage, a setting that motivates practical algorithms such as Conservative -Learning (CQL; Ku…
An Improved Algorithm for Adversarial Linear Contextual Bandits via Reduction
Tim van Erven, Jack Mayo, Julia Olkhovskaya +1
We present an oracle-efficient, near-optimal algorithm for linear contextual bandits with adversarial losses and stochastic action sets, only requiring a linear optimization oracle…
An Improved Model-Free Decision-Estimation Coefficient with Applications in Adversarial MDPs
Haolin Liu, Chen-Yu Wei, Julian Zimmert
We study decision making with structured observation (DMSO). Previous work (Foster et al., 2021b, 2023a) has characterized the complexity of DMSO via the decision-estimation coeffi…
Achieving Optimal Static and Dynamic Regret Simultaneously in Bandits with Deterministic Losses
Jian Qian, Chen-Yu Wei
In adversarial multi-armed bandits, two performance measures are commonly used: static regret, which compares the learner to the best fixed arm, and dynamic regret, which compares…
Is Online Linear Optimization Sufficient for Strategic Robustness?
Yang Cai, Haipeng Luo, Chen-Yu Wei +1
We consider bidding in repeated Bayesian first-price auctions. Bidding algorithms that achieve optimal regret have been extensively studied, but their strategic robustness to the s…
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
Haolin Liu, Dian Yu, Sidi Lu +6
Reinforcement learning (RL) has emerged as a powerful framework for improving the reasoning capabilities of large language models (LLMs). However, most existing RL approaches rely…