Showing cs.LGShow all
3 papers · 1 filter
cs.LG2019
Revisit Policy Optimization in Matrix Form
Sitao Luan, Xiao-Wen Chang, Doina Precup
In tabular case, when the reward and environment dynamics are known, policy evaluation can be written as , where is t…
cs.LG2019
Break the Ceiling: Stronger Multi-scale Deep Graph Convolutional Networks
Sitao Luan, Mingde Zhao, Xiao-Wen Chang +1
Recently, neural network based approaches have achieved significant improvement for solving large, complex, graph-structured problems. However, their bottlenecks still need to be a…
cs.LG2019
META-Learning State-based Eligibility Traces for More Sample-Efficient Policy Evaluation
Mingde Zhao, Sitao Luan, Ian Porada +2
Temporal-Difference (TD) learning is a standard and very successful reinforcement learning approach, at the core of both algorithms that learn the value of a given policy, as well…