4 papers
An MRP Formulation for Supervised Learning: Generalized Temporal Difference Learning Models
Yangchen Pan, Junfeng Wen, Chenjun Xiao +1
In traditional statistical learning, data points are usually assumed to be independently and identically distributed (i.i.d.) following an unknown probability distribution. This pa…
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning
Chen-Xiao Gao, Chenyang Wu, Mingjun Cao +3
Behavior regularization, which constrains the policy to stay close to some behavior policy, is widely used in offline reinforcement learning (RL) to manage the risk of hazardous ex…
Large Language Model-Enhanced Multi-Armed Bandits
Jiahang Sun, Zhiyong Wang, Runhan Yang +3
Large language models (LLMs) have been adopted to solve sequential decision-making tasks such as multi-armed bandits (MAB), in which an LLM is directly instructed to select the arm…
Diffusion Spectral Representation for Reinforcement Learning
Dmitry Shribak, Chen-Xiao Gao, Yitong Li +2
Diffusion-based models have achieved notable empirical successes in reinforcement learning (RL) due to their expressiveness in modeling complex distributions. Despite existing meth…