10 papers
Constrained Online Convex Optimization without Slater's Condition
Kihyun Yu, Junehee Lee, Dabeen Lee
We study constrained online convex optimization with adversarial losses and stochastic or adversarial constraints. For stochastic constraints, existing algorithms that achieve near…
Learning to Route and Schedule LLMs from User Retrials via Contextual Queueing Bandits
Seoungbin Bae, Junyoung Son, Dabeen Lee
Explosive demands for LLMs often cause user queries to accumulate in server queues, requiring efficient routing (query-LLM matching) and scheduling (query prioritization) mechanism…
Algorithm for Contextual Queueing Bandits with Rate-Optimal Queue Length Regret
Seoungbin Bae, Dabeen Lee
Contextual queueing bandits provide a framework for learning to schedule heterogeneous jobs under unknown context-dependent service rates. Under stochastic contexts, existing algor…
Neural Logistic Bandits
Seoungbin Bae, Dabeen Lee
We study the problem of neural logistic bandits, where the main task is to learn an unknown reward function within a logistic link function using a neural network. Existing approac…
Queue Length Regret Bounds for Contextual Queueing Bandits
Seoungbin Bae, Garyeong Kang, Dabeen Lee
We introduce contextual queueing bandits, a new context-aware framework for scheduling while simultaneously learning unknown service rates. Individual jobs carry heterogeneous cont…
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
Kihyun Yu, Beomhan Baek, Dabeen Lee
We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Our contributions are twofold. First, we establish…