5 citations · 5 across the 4 of their papers we have counts for
7 papers
Off-Policy Risk Assessment in Markov Decision Processes
Audrey Huang, Liu Leqi, Zachary Chase Lipton +1
Addressing such diverse ends as safety alignment with human preferences, and the efficiency of learning, a growing line of reinforcement learning research focuses on risk functiona…
When Curation Becomes Creation: Algorithms, Microcontent, and the Vanishing Distinction between Platforms and Creators
Liu Leqi, Dylan Hadfield-Menell, Zachary C. Lipton
Ever since social activity on the Internet began migrating from the wilds of the open web to the walled gardens erected by so-called platforms, debates have raged about the respons…
Off-Policy Risk Assessment in Contextual Bandits
Audrey Huang, Liu Leqi, Zachary C. Lipton +1
Even when unable to run experiments, practitioners can evaluate prospective policies, using previously logged data. However, while the bandits literature has adopted a diverse set…
On the Convergence and Optimality of Policy Gradient for Markov Coherent Risk
Audrey Huang, Liu Leqi, Zachary C. Lipton +1
In order to model risk aversion in reinforcement learning, an emerging line of research adapts familiar algorithms to optimize coherent risk functionals, a class that includes cond…
Rebounding Bandits for Modeling Satiation Effects
Liu Leqi, Fatma Kilinc-Karzan, Zachary C. Lipton +1
Psychological research shows that enjoyment of many goods is subject to satiation, with short-term satisfaction declining after repeated exposures to the same item. Nevertheless, p…
Game Design for Eliciting Distinguishable Behavior
Fan Yang, Liu Leqi, Yifan Wu +4
The ability to inferring latent psychological traits from human behavior is key to developing personalized human-interacting machine learning systems. Approaches to infer such trai…