4 papers
Bilevel Optimization over Saddle Points of Zero-Sum Markov Games
Zihao Zheng, Irwin King, Songtao Lu
Reinforcement learning (RL) often has a hierarchical structure, where an upper-level (UL) learner selects model parameters and a lower-level (LL) decision-making process responds,…
Diversity or Precision? A Deep Dive into Next Token Prediction
Haoyuan Wu, Hai Wang, Jiajia Wu +5
Recent advancements have shown that reinforcement learning (RL) can substantially improve the reasoning abilities of large language models (LLMs). The effectiveness of such RL trai…
Reinforcement Learning on Pre-Training Data
Siheng Li, Kejiao Li, Zenan Xu +33
The growing disparity between the exponential scaling of computational resources and the finite growth of high-quality text data now constrains conventional scaling approaches for…
RaSeRec: Retrieval-Augmented Sequential Recommendation
Xinping Zhao, Baotian Hu, Yan Zhong +5
Although prevailing supervised and self-supervised learning augmented sequential recommendation (SeRec) models have achieved improved performance with powerful neural network archi…