22 citations · 39 across the 23 of their papers we have counts for
21 papers · 1 filter
A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics
Zheshun Wu, Renjie Zheng, Jinhang Zuo +2
This paper investigates a hybrid reinforcement learning setting in tabular Markov Decision Processes (MDPs), where an agent aims to learn an optimal policy by combining online inte…
Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation
Xutong Liu, Baran Atalar, Xiangxiang Dai +5
Large Language Models (LLMs) are revolutionizing how users interact with information systems, yet their high inference cost poses serious scalability and sustainability challenges.…
Online Multi-LLM Selection via Contextual Bandits under Unstructured Context Evolution
Manhin Poon, XiangXiang Dai, Xutong Liu +3
Large language models (LLMs) exhibit diverse response behaviors, costs, and strengths, making it challenging to select the most suitable LLM for a given user query. We study the pr…
A Unified Online-Offline Framework for Co-Branding Campaign Recommendations
Xiangxiang Dai, Xiaowei Sun, Jinhang Zuo +2
Co-branding has become a vital strategy for businesses aiming to expand market reach within recommendation systems. However, identifying effective cross-industry partnerships remai…
Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection
Qirun Zeng, Eric He, Richard Hoffmann +2
Adversarial attacks on stochastic bandits have traditionally relied on some unrealistic assumptions, such as per-round reward manipulation and unbounded perturbations, limiting the…
Fusing Reward and Dueling Feedback in Stochastic Bandits
Xuchuang Wang, Qirun Zeng, Jinhang Zuo +4
This paper investigates the fusion of absolute (reward) and relative (dueling) feedback in stochastic bandits, where both feedback types are gathered in each decision round. We der…