5 citations · 6 across the 4 of their papers we have counts for
4 papers
Multiple Policy Value Monte Carlo Tree Search
Li-Cheng Lan, Wei Li, Ting-Han Wei +1
Many of the strongest game playing programs use a combination of Monte Carlo tree search (MCTS) and deep neural networks (DNN), where the DNNs are used as policy or value evaluator…
Towards Combining On-Off-Policy Methods for Real-World Applications
Kai-Chun Hu, Chen-Huan Pi, Ting Han Wei +4
In this paper, we point out a fundamental property of the objective in reinforcement learning, with which we can reformulate the policy gradient objective into a perceptron-like lo…
Comparison Training for Computer Chinese Chess
Wen-Jie Tseng, Jr-Chang Chen, I-Chen Wu +1
This paper describes the application of comparison training (CT) for automatic feature weight tuning, with the final objective of improving the evaluation functions used in Chinese…
Multi-Labelled Value Networks for Computer Go
Ti-Rong Wu, I-Chen Wu, Guan-Wun Chen +4
This paper proposes a new approach to a novel value network architecture for the game Go, called a multi-labelled (ML) value network. In the ML value network, different values (win…