3 papers
cs.LG2026
Gradient Iterated Temporal-Difference Learning
Théo Vincent, Kevin Gerhardt, Yogesh Tripathi +5
Temporal-difference (TD) learning is highly effective at controlling and evaluating an agent's long-term outcomes. Most approaches in this paradigm implement a semi-gradient update…
cs.LG2026
Bridging the Performance Gap Between Target-Free and Target-Based Reinforcement Learning
Théo Vincent, Yogesh Tripathi, Tim Faust +5
The use of target networks in deep reinforcement learning is a widely popular solution to mitigate the brittleness of semi-gradient approaches and stabilize learning. However, targ…
cs.LG2025
Eau De -Network: Adaptive Distillation of Neural Networks in Deep Reinforcement Learning
Théo Vincent, Tim Faust, Yogesh Tripathi +2
Recent works have successfully demonstrated that sparse deep reinforcement learning agents can be competitive against their dense counterparts. This opens up opportunities for rein…