2 papers
cs.CL2019
Reducing Sentiment Bias in Language Models via Counterfactual Evaluation
Po-Sen Huang, Huan Zhang, Ray Jiang +6
Advances in language modeling architectures and the availability of large text corpora have driven progress in automatic text generation. While this results in models capable of ge…
cs.LG2018
Scalable agent alignment via reward modeling: a research direction
Jan Leike, David Krueger, Tom Everitt +3
One obstacle to applying reinforcement learning algorithms to real-world problems is the lack of suitable reward functions. Designing such reward functions is difficult in part bec…