5.7k citations · 6.1k across the 5 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2016★ 5 cited
Improving Policy Gradient by Exploring Under-appreciated Rewards
Ofir Nachum, Mohammad Norouzi, Dale Schuurmans
This paper presents a novel form of policy gradient for model-free reinforcement learning (RL) with improved exploration properties. Current policy-based methods use entropy regula…
cs.LG2016★ 88 cited
Reward Augmented Maximum Likelihood for Neural Structured Prediction
Mohammad Norouzi, Samy Bengio, Zhifeng Chen +4
A key problem in structured output prediction is direct optimization of the task reward function that matters for test evaluation. This paper presents a simple and computationally…