18 citations · 28 across the 44 of their papers we have counts for
Showing 2021Show all
2 papers · 1 filter
cs.LG2021★ 1 cited
Balanced Q-learning: Combining the Influence of Optimistic and Pessimistic Targets
Thommen George Karimpanal, Hung Le, Majid Abdolshah +4
The optimistic nature of the Q-learning target leads to an overestimation bias, which is an inherent problem associated with standard learning. Such a bias fails to account for…
cs.LG2021★ 2 cited
Offline Neural Contextual Bandits: Pessimism, Optimization and Generalization
Thanh Nguyen-Tang, Sunil Gupta, A. Tuan Nguyen +1
Offline policy learning (OPL) leverages existing data collected a priori for policy optimization without any active exploration. Despite the prevalence and recent interest in this…