Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning
Ahmed Hendawy, Henrik Metternich, Théo Vincent +3
The use of target networks is a popular approach for estimating value functions in deep Reinforcement Learning (RL). While effective, the target network remains a compromise soluti…
cs.LG2026
Do Not Imitate, Reinforce: Iterative Classification via Belief Refinement
Mahdi Kallel, Johannes Tölle, Ahmed Hendawy +1
Standard supervised classification trains models to imitate the exact labels provided by a perfect oracle. This imitation happens in a single pass, restricting the model to a fixed…
cs.LG2024
Augmented Bayesian Policy Search
Mahdi Kallel, Debabrota Basu, Riad Akrour +1
Deterministic policies are often preferred over stochastic ones when implemented on physical systems. They can prevent erratic and harmful behaviors while being easier to implement…