2 papers
cs.LG2026
DART: aDaptive Accept RejecT for non-linear top-K subset identification
Mridul Agarwal, Vaneet Aggarwal, Christopher J. Quinn +1
We consider the bandit problem of selecting out of arms at each time step. The reward can be a non-linear function of the rewards of the selected individual arms. The direc…
cs.LG2025
Joint Optimization of Multi-Objective Reinforcement Learning with Policy Gradient Based Algorithm
Qinbo Bai, Mridul Agarwal, Vaneet Aggarwal
Many engineering problems have multiple objectives, and the overall aim is to optimize a non-linear function of these objectives. In this paper, we formulate the problem of maximiz…