5 citations · 10 across the 6 of their papers we have counts for
Showing 2019Show all
2 papers · 1 filter
cs.AI2019
Multiple Policy Value Monte Carlo Tree Search
Li-Cheng Lan, Wei Li, Ting-Han Wei +1
Many of the strongest game playing programs use a combination of Monte Carlo tree search (MCTS) and deep neural networks (DNN), where the DNNs are used as policy or value evaluator…
cs.LG2019★ 1 cited
Towards Combining On-Off-Policy Methods for Real-World Applications
Kai-Chun Hu, Chen-Huan Pi, Ting Han Wei +4
In this paper, we point out a fundamental property of the objective in reinforcement learning, with which we can reformulate the policy gradient objective into a perceptron-like lo…