Publications (12)
Beyond Greedy Ranking: Slate Optimization via List-CVAE
Ray Jiang, Sven Gowal, Timothy A. Mann +1
The conventional solution to the recommendation problem greedily ranks individual document candidates by prediction scores. However, this method fails to optimize the slate as a wh…
Reducing Sentiment Bias in Language Models via Counterfactual Evaluation
Po-Sen Huang, Huan Zhang, Ray Jiang +6
Advances in language modeling architectures and the availability of large text corpora have driven progress in automatic text generation. While this results in models capable of ge…
Learning from Delayed Outcomes via Proxies with Applications to Recommender Systems
Timothy A. Mann, Sven Gowal, András György +4
Predicting delayed outcomes is an important problem in recommender systems (e.g., if customers will finish reading an ebook). We formalize the problem as an adversarial, delayed on…
Causally Correct Partial Models for Reinforcement Learning
Danilo J. Rezende, Ivo Danihelka, George Papamakarios +11
In reinforcement learning, we can learn a model of future observations and rewards, and use it to plan the agent's next actions. However, jointly modeling future observations can b…
Wasserstein Fair Classification
Ray Jiang, Aldo Pacchiano, Tom Stepleton +2
We propose an approach to fair classification that enforces independence between the classifier outputs and sensitive information by minimizing Wasserstein-1 distances. The approac…
AlphaStar Unplugged: Large-Scale Offline Reinforcement Learning
Michaël Mathieu, Sherjil Ozair, Srivatsan Srinivasan +21
StarCraft II is one of the most challenging simulated reinforcement learning environments; it is partially observable, stochastic, multi-agent, and mastering StarCraft II requires…