11 papers
A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics
Zheshun Wu, Renjie Zheng, Jinhang Zuo +2
This paper investigates a hybrid reinforcement learning setting in tabular Markov Decision Processes (MDPs), where an agent aims to learn an optimal policy by combining online inte…
Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback
Qirun Zeng, Xuchuang Wang, Jiayi Shen +3
We study fixed-confidence best arm identification in generalized linear bandits under a hybrid feedback model: at each round, the learner may query either (i) absolute reward feedb…
Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection
Qirun Zeng, Eric He, Richard Hoffmann +2
Adversarial attacks on stochastic bandits have traditionally relied on some unrealistic assumptions, such as per-round reward manipulation and unbounded perturbations, limiting the…
Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation
Xutong Liu, Baran Atalar, Xiangxiang Dai +5
Large Language Models (LLMs) are revolutionizing how users interact with information systems, yet their high inference cost poses serious scalability and sustainability challenges.…
Online Multi-LLM Selection via Contextual Bandits under Unstructured Context Evolution
Manhin Poon, XiangXiang Dai, Xutong Liu +3
Large language models (LLMs) exhibit diverse response behaviors, costs, and strengths, making it challenging to select the most suitable LLM for a given user query. We study the pr…
Offline Learning for Combinatorial Multi-armed Bandits
Xutong Liu, Xiangxiang Dai, Jinhang Zuo +4
The combinatorial multi-armed bandit (CMAB) is a fundamental sequential decision-making framework, extensively studied over the past decade. However, existing work primarily focuse…