2 papers
cs.LG2025
Large Language Model-Enhanced Multi-Armed Bandits
Jiahang Sun, Zhiyong Wang, Runhan Yang +3
Large language models (LLMs) have been adopted to solve sequential decision-making tasks such as multi-armed bandits (MAB), in which an LLM is directly instructed to select the arm…
cs.LG2025
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning
Chen-Xiao Gao, Chenyang Wu, Mingjun Cao +3
Behavior regularization, which constrains the policy to stay close to some behavior policy, is widely used in offline reinforcement learning (RL) to manage the risk of hazardous ex…