1 paper
Haoyu Wei, Runzhe Wan, Lei Shi +1
Many real-world bandit applications are characterized by sparse rewards, which can significantly hinder learning efficiency. Leveraging problem-specific structures for careful dist…