collaborators

11 papers

cs.LG2026

A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics

Zheshun Wu, Renjie Zheng, Jinhang Zuo +2

This paper investigates a hybrid reinforcement learning setting in tabular Markov Decision Processes (MDPs), where an agent aims to learn an optimal policy by combining online inte…

cs.AI2026

Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback

Qirun Zeng, Xuchuang Wang, Jiayi Shen +3

We study fixed-confidence best arm identification in generalized linear bandits under a hybrid feedback model: at each round, the learner may query either (i) absolute reward feedb…

cs.LG2026

Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection

Qirun Zeng, Eric He, Richard Hoffmann +2

Adversarial attacks on stochastic bandits have traditionally relied on some unrealistic assumptions, such as per-round reward manipulation and unbounded perturbations, limiting the…

cs.LG2026

Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation

Xutong Liu, Baran Atalar, Xiangxiang Dai +5

Large Language Models (LLMs) are revolutionizing how users interact with information systems, yet their high inference cost poses serious scalability and sustainability challenges.…

cs.LG2025

Online Multi-LLM Selection via Contextual Bandits under Unstructured Context Evolution

Manhin Poon, XiangXiang Dai, Xutong Liu +3

Large language models (LLMs) exhibit diverse response behaviors, costs, and strengths, making it challenging to select the most suitable LLM for a given user query. We study the pr…

cs.LG2025

Offline Learning for Combinatorial Multi-armed Bandits

Xutong Liu, Xiangxiang Dai, Jinhang Zuo +4

The combinatorial multi-armed bandit (CMAB) is a fundamental sequential decision-making framework, extensively studied over the past decade. However, existing work primarily focuse…