2 papers
cs.LG2025
Reinforcement Learning From State and Temporal Differences
Lex Weaver, Jonathan Baxter
TD() with function approximation has proved empirically successful for some complex reinforcement learning problems. For linear approximation, TD() has been shown to minimi…
cs.LG2025
A Multi-Agent, Policy-Gradient approach to Network Routing
Nigel Tao, Jonathan Baxter, Lex Weaver
Network routing is a distributed decision problem which naturally admits numerical performance measures, such as the average time for a packet to travel from source to destination.…