5 papers
Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
Rui Hu, Yu Chen, Longbo Huang
Many popular practical reinforcement learning (RL) algorithms employ evolving reward functions-through techniques such as reward shaping, entropy regularization, or curriculum lear…
Finite-Time Convergence Analysis of ODE-based Generative Models for Stochastic Interpolants
Yuhao Liu, Rui Hu, Yu Chen +1
Stochastic interpolants offer a robust framework for continuously transforming samples between arbitrary data distributions, holding significant promise for generative modeling. De…
Finite-Time Analysis of Discrete-Time Stochastic Interpolants
Yuhao Liu, Yu Chen, Rui Hu +1
The stochastic interpolant framework offers a powerful approach for constructing generative models based on ordinary differential equations (ODEs) or stochastic differential equati…
uniINF: Best-of-Both-Worlds Algorithm for Parameter-Free Heavy-Tailed MABs
Yu Chen, Jiatai Huang, Yan Dai +1
In this paper, we present a novel algorithm, uniINF, for the Heavy-Tailed Multi-Armed Bandits (HTMAB) problem, demonstrating robustness and adaptability in both stochastic and adve…
Adversarial Network Optimization under Bandit Feedback: Maximizing Utility in Non-Stationary Multi-Hop Networks
Yan Dai, Longbo Huang
Stochastic Network Optimization (SNO) concerns scheduling in stochastic queueing systems. It has been widely studied in network theory. Classical SNO algorithms require network con…