1 citations · 2 across the 4 of their papers we have counts for
4 papers
Nonparametric Learning and Earning with One-Point Feedback under Nonstationarity
Xiangyu Yang, Feng Xu, Jian-Qiang Hu +1
Firms increasingly rely on dynamic pricing to respond to evolving customer demand, yet in many applications they observe only the revenue generated by a single posted price in each…
Stochastic Approximation Methods for Distortion Risk Measure Optimization
Jinyang Jiang, Bernd Heidergott, Jiaqiao Hu +1
Distortion Risk Measures (DRMs) capture risk preferences in decision-making and serve as general criteria for managing uncertainty. This paper proposes gradient descent algorithms…
Quantile-Based Deep Reinforcement Learning using Two-Timescale Policy Gradient Algorithms
Jinyang Jiang, Jiaqiao Hu, Yijie Peng
Classical reinforcement learning (RL) aims to optimize the expected cumulative reward. In this work, we consider the RL setting where the goal is to optimize the quantile of the cu…
Quantile-Based Policy Optimization for Reinforcement Learning
Jinyang Jiang, Jiaqiao Hu, Yijie Peng
Classical reinforcement learning (RL) aims to optimize the expected cumulative rewards. In this work, we consider the RL setting where the goal is to optimize the quantile of the c…