1 paper
Mengfan Xu, Diego Klabjan
We study the challenging exploration incentive problem in both bandit and reinforcement learning, where the rewards are scale-free and potentially unbounded, driven by real-world s…