5 citations · 5 across the 3 of their papers we have counts for
6 papers
Beyond the Return: Off-policy Function Estimation under User-specified Error-measuring Distributions
Audrey Huang, Nan Jiang
Off-policy evaluation often refers to two related tasks: estimating the expected return of a policy and estimating its value function (or other functions of interest, such as densi…
Off-Policy Risk Assessment in Markov Decision Processes
Audrey Huang, Liu Leqi, Zachary Chase Lipton +1
Addressing such diverse ends as safety alignment with human preferences, and the efficiency of learning, a growing line of reinforcement learning research focuses on risk functiona…
Off-Policy Risk Assessment in Contextual Bandits
Audrey Huang, Liu Leqi, Zachary C. Lipton +1
Even when unable to run experiments, practitioners can evaluate prospective policies, using previously logged data. However, while the bandits literature has adopted a diverse set…
On the Convergence and Optimality of Policy Gradient for Markov Coherent Risk
Audrey Huang, Liu Leqi, Zachary C. Lipton +1
In order to model risk aversion in reinforcement learning, an emerging line of research adapts familiar algorithms to optimize coherent risk functionals, a class that includes cond…
Graph-Structured Visual Imitation
Maximilian Sieb, Zhou Xian, Audrey Huang +2
We cast visual imitation as a visual correspondence problem. Our robotic agent is rewarded when its actions result in better matching of relative spatial configurations for corresp…
Physically optimizing inference
Audrey Huang, Benjamin Sheldan, David A. Sivak +1
Data is scaling exponentially in fields ranging from genomics to neuroscience to economics. A central question is: can modern machine learning methods be applied to construct predi…