3 papers
cs.LG2025
Capacity-Constrained Continual Learning
Zheng Wen, Doina Precup, Benjamin Van Roy +1
Any agents we can possibly build are subject to capacity constraints, as memory and compute resources are inherently finite. However, comparatively little attention has been dedica…
cs.LG2024
RLHF and IIA: Perverse Incentives
Wanqiao Xu, Shi Dong, Xiuyuan Lu +3
Existing algorithms for reinforcement learning from human feedback (RLHF) can incentivize responses at odds with preferences because they are based on models that assume independen…
cs.IR2023
Design Principles of Robust Multi-Armed Bandit Framework in Video Recommendations
Belhassen Bayar, Phanideep Gampa, Ainur Yessenalina +1
Current multi-armed bandit approaches in recommender systems (RS) have focused more on devising effective exploration techniques, while not adequately addressing common exploitatio…