31 citations · 31 across the 1 of their papers we have counts for
2 papers
cs.AI2017★ 31 cited
End-to-End Offline Goal-Oriented Dialog Policy Learning via Policy Gradient
Li Zhou, Kevin Small, Oleg Rokhlenko +1
Learning a goal-oriented dialog policy is generally performed offline with supervised learning algorithms or online with reinforcement learning (RL). Additionally, as companies acc…
cs.LG2016
Latent Contextual Bandits and their Application to Personalized Recommendations for New Users
Li Zhou, Emma Brunskill
Personalized recommendations for new users, also known as the cold-start problem, can be formulated as a contextual bandit problem. Existing contextual bandit algorithms generally…