3 papers
cs.LG2026
Algorithm for Contextual Queueing Bandits with Rate-Optimal Queue Length Regret
Seoungbin Bae, Dabeen Lee
Contextual queueing bandits provide a framework for learning to schedule heterogeneous jobs under unknown context-dependent service rates. Under stochastic contexts, existing algor…
cs.LG2026
Logistic Bandits with Regret without Context Diversity Assumptions
Seoungbin Bae, Dabeen Lee
We study the -armed logistic bandit problem, where at each round, the agent observes feature vectors associated with actions. Existing approaches that achieve a rate-opt…
cs.LG2026
Near-Optimal Primal-Dual Algorithm for Learning Linear Mixture CMDPs with Adversarial Rewards
Kihyun Yu, Seoungbin Bae, Dabeen Lee
We study safe reinforcement learning in finite-horizon linear mixture constrained Markov decision processes (CMDPs) with adversarial rewards under full-information feedback and an…