24 citations · 114 across the 29 of their papers we have counts for
13 papers · 1 filter
Beyond Black-Box Advice: Learning-Augmented Algorithms for MDPs with Q-Value Predictions
Tongxin Li, Yiheng Lin, Shaolei Ren +1
We study the tradeoff between consistency and robustness in the context of a single-trajectory time-varying Markov Decision Process (MDP) with untrusted machine-learned advice. Our…
Adversarial Attacks on Online Learning to Rank with Click Feedback
Jinhang Zuo, Zhiyao Zhang, Zhiyong Wang +3
Online learning to rank (OLTR) is a sequential decision-making problem where a learning agent selects an ordered list of items and receives feedback through user clicks. Although p…
Contextual Combinatorial Bandits with Probabilistically Triggered Arms
Xutong Liu, Jinhang Zuo, Siwei Wang +4
We study contextual combinatorial bandits with probabilistically triggered arms (CMAB-T) under a variety of smoothness conditions that capture a wide range of applications, suc…
Convergence Rates for Localized Actor-Critic in Networked Markov Potential Games
Zhaoyi Zhou, Zaiwei Chen, Yiheng Lin +1
We introduce a class of networked Markov potential games in which agents are associated with nodes in a network. Each agent has its own local potential function, and the reward of…
Global Convergence of Localized Policy Iteration in Networked Multi-Agent Reinforcement Learning
Yizhou Zhang, Guannan Qu, Pan Xu +3
We study a multi-agent reinforcement learning (MARL) problem where the agents interact over a given network. The goal of the agents is to cooperatively maximize the average of thei…
Online Optimization with Feedback Delay and Nonlinear Switching Cost
Weici Pan, Guanya Shi, Yiheng Lin +1
We study a variant of online optimization in which the learner receives -round about hitting cost and there is a multi-step nonlinear switching cost,…