adaptive discretization 1combinatorial bandits 1confidence bounds 1continuous outcomes 1sublinear regret 1
From the 1 of 1 linked paper with an AI index.
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
On the Sublinear Regret of Continuous K-Max Bandits
Yu Chen, Siwei Wang, Longbo Huang +1
The paper studies continuous K‑max combinatorial multi‑armed bandits, proposing the DCK‑UCB algorithm with adaptive discretization and bias‑corrected confidence bounds that achieve…
cs.LG2024
Provable Risk-Sensitive Distributional Reinforcement Learning with General Function Approximation
Yu Chen, Xiangcheng Zhang, Siwei Wang +1
In the realm of reinforcement learning (RL), accounting for risk is crucial for making decisions under uncertainty, particularly in applications where safety and reliability are pa…