From the 1 of 8 linked papers with an AI index.
8 papers
On the Sublinear Regret of Continuous K-Max Bandits
Yu Chen, Siwei Wang, Longbo Huang +1
The paper studies continuous K‑max combinatorial multi‑armed bandits, proposing the DCK‑UCB algorithm with adaptive discretization and bias‑corrected confidence bounds that achieve…
When Context Returns: Toward Robust Internalization in On-Policy Distillation
Xun Wang, Ruishuo Chen, Zhuoran Li +2
Recent work has shown that on-policy distillation can internalize privileged context, such as system prompts or task hints, into a student model so that the context is no longer ne…
Best-of-Both-Worlds for Heavy-Tailed Markov Decision Processes
Yu Chen, Yuhao Liu, Jiatai Huang +2
We investigate episodic Markov Decision Processes with heavy-tailed losses (HTMDPs). Existing approaches for HTMDPs are conservative in stochastic environments and lack adaptivity…
Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
Rui Hu, Yu Chen, Longbo Huang
Many popular practical reinforcement learning (RL) algorithms employ evolving reward functions-through techniques such as reward shaping, entropy regularization, or curriculum lear…
Finite-Time Convergence Analysis of ODE-based Generative Models for Stochastic Interpolants
Yuhao Liu, Rui Hu, Yu Chen +1
Stochastic interpolants offer a robust framework for continuously transforming samples between arbitrary data distributions, holding significant promise for generative modeling. De…
Finite-Time Analysis of Discrete-Time Stochastic Interpolants
Yuhao Liu, Yu Chen, Rui Hu +1
The stochastic interpolant framework offers a powerful approach for constructing generative models based on ordinary differential equations (ODEs) or stochastic differential equati…