5 citations · 9 across the 5 of their papers we have counts for
6 papers · 1 filter
Reinforcement Learning in Low-Rank MDPs with Density Features
Audrey Huang, Jinglin Chen, Nan Jiang
MDPs with low-rank transitions -- that is, the transition matrix can be factored into the product of two matrices, left and right -- is a highly representative structure that enabl…
Beyond the Return: Off-policy Function Estimation under User-specified Error-measuring Distributions
Audrey Huang, Nan Jiang
Off-policy evaluation often refers to two related tasks: estimating the expected return of a policy and estimating its value function (or other functions of interest, such as densi…
Off-Policy Risk Assessment in Markov Decision Processes
Audrey Huang, Liu Leqi, Zachary Chase Lipton +1
Addressing such diverse ends as safety alignment with human preferences, and the efficiency of learning, a growing line of reinforcement learning research focuses on risk functiona…
Offline Reinforcement Learning with Realizability and Single-policy Concentrability
Wenhao Zhan, Baihe Huang, Audrey Huang +2
Sample-efficiency guarantees for offline reinforcement learning (RL) often rely on strong assumptions on both the function classes (e.g., Bellman-completeness) and the data coverag…
Off-Policy Risk Assessment in Contextual Bandits
Audrey Huang, Liu Leqi, Zachary C. Lipton +1
Even when unable to run experiments, practitioners can evaluate prospective policies, using previously logged data. However, while the bandits literature has adopted a diverse set…
On the Convergence and Optimality of Policy Gradient for Markov Coherent Risk
Audrey Huang, Liu Leqi, Zachary C. Lipton +1
In order to model risk aversion in reinforcement learning, an emerging line of research adapts familiar algorithms to optimize coherent risk functionals, a class that includes cond…