activity
20192026
most citedThe Online Pause and Resume Problem: Optimal Algorithms and An Application to Carbon-Aware Load Shifting

22 citations · 39 across the 23 of their papers we have counts for

collaborators
Showing cs.LGShow all

21 papers · 1 filter

cs.LG2026

A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics

Zheshun Wu, Renjie Zheng, Jinhang Zuo +2

This paper investigates a hybrid reinforcement learning setting in tabular Markov Decision Processes (MDPs), where an agent aims to learn an optimal policy by combining online inte…

cs.LG2025

Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation

Xutong Liu, Baran Atalar, Xiangxiang Dai +5

Large Language Models (LLMs) are revolutionizing how users interact with information systems, yet their high inference cost poses serious scalability and sustainability challenges.…

cs.LG2025

Online Multi-LLM Selection via Contextual Bandits under Unstructured Context Evolution

Manhin Poon, XiangXiang Dai, Xutong Liu +3

Large language models (LLMs) exhibit diverse response behaviors, costs, and strengths, making it challenging to select the most suitable LLM for a given user query. We study the pr…

cs.LG2025★ 2 cited

A Unified Online-Offline Framework for Co-Branding Campaign Recommendations

Xiangxiang Dai, Xiaowei Sun, Jinhang Zuo +2

Co-branding has become a vital strategy for businesses aiming to expand market reach within recommendation systems. However, identifying effective cross-industry partnerships remai…

cs.LG2025

Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection

Qirun Zeng, Eric He, Richard Hoffmann +2

Adversarial attacks on stochastic bandits have traditionally relied on some unrealistic assumptions, such as per-round reward manipulation and unbounded perturbations, limiting the…

cs.LG2025

Fusing Reward and Dueling Feedback in Stochastic Bandits

Xuchuang Wang, Qirun Zeng, Jinhang Zuo +4

This paper investigates the fusion of absolute (reward) and relative (dueling) feedback in stochastic bandits, where both feedback types are gathered in each decision round. We der…