collaborators

6 papers

cs.LG2026

A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics

Zheshun Wu, Renjie Zheng, Jinhang Zuo +2

This paper investigates a hybrid reinforcement learning setting in tabular Markov Decision Processes (MDPs), where an agent aims to learn an optimal policy by combining online inte…

cs.LG2026

Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift

Bochao Li, Yao Fu, Wei Chen +1

Offline-to-online learning aims to improve online decision-making by leveraging offline logged data. A central challenge in this setting is the distribution shift between offline a…

stat.ML2026

A Regret Perspective on Online Multiple Testing

Qingyang Hao, Kongchang Zhou, Fang Kong +1

Online Multiple Testing (OMT), a fundamental pillar of sequential statistical inference, traditionally evaluates the False Discovery Rate (FDR) and statistical power in isolation,…

cs.AI2026

Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback

Qirun Zeng, Xuchuang Wang, Jiayi Shen +3

We study fixed-confidence best arm identification in generalized linear bandits under a hybrid feedback model: at each round, the learner may query either (i) absolute reward feedb…

cs.LG2025

Hybrid Combinatorial Multi-armed Bandits with Probabilistically Triggered Arms

Kongchang Zhou, Tingyu Zhang, Wei Chen +1

The problem of combinatorial multi-armed bandits with probabilistically triggered arms (CMAB-T) has been extensively studied. Prior work primarily focuses on either the online sett…

cs.LG2025

Online Multi-LLM Selection via Contextual Bandits under Unstructured Context Evolution

Manhin Poon, XiangXiang Dai, Xutong Liu +3

Large language models (LLMs) exhibit diverse response behaviors, costs, and strengths, making it challenging to select the most suitable LLM for a given user query. We study the pr…