4 papers · 1 filter
Revisit Policy Optimization in Matrix Form
Sitao Luan, Xiao-Wen Chang, Doina Precup
In tabular case, when the reward and environment dynamics are known, policy evaluation can be written as , where is t…
Break the Ceiling: Stronger Multi-scale Deep Graph Convolutional Networks
Sitao Luan, Mingde Zhao, Xiao-Wen Chang +1
Recently, neural network based approaches have achieved significant improvement for solving large, complex, graph-structured problems. However, their bottlenecks still need to be a…
Improved Upper Bounds on the Hermite and KZ Constants
Jinming Wen, Xiao-Wen Chang, Jian Weng
The Korkine-Zolotareff (KZ) reduction is a widely used lattice reduction strategy in communications and cryptography. The Hermite constant, which is a vital constant of lattice, ha…
META-Learning State-based Eligibility Traces for More Sample-Efficient Policy Evaluation
Mingde Zhao, Sitao Luan, Ian Porada +2
Temporal-Difference (TD) learning is a standard and very successful reinforcement learning approach, at the core of both algorithms that learn the value of a given policy, as well…