2 papers
cs.LG2024
Pearl: A Production-ready Reinforcement Learning Agent
Zheqing Zhu, Rodrigo de Salvo Braz, Jalaj Bhandari +12
Reinforcement learning (RL) is a versatile framework for optimizing long-term goals. Although many real-world problems can be formalized with RL, learning and deploying a performan…
cs.LG2024
Uncertainty of Joint Neural Contextual Bandit
Hongbo Guo, Zheqing Zhu
Contextual bandit learning is increasingly favored in modern large-scale recommendation systems. To better utlize the contextual information and available user or item features, th…