2 papers
cs.LG2026
Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
Chaorui Yao, Yanxi Chen, Yuchang Sun +5
Off-policy reinforcement learning (RL) for large language models (LLMs) is attracting growing interest, driven by practical constraints in real-world applications, the complexity o…
cs.AI2025
Instructional Prompt Optimization for Few-Shot LLM-Based Recommendations on Cold-Start Users
Haowei Yang, Yushang Zhao, Sitao Min +3
The cold-start user issue further compromises the effectiveness of recommender systems in limiting access to the historical behavioral information. It is an effective pipeline to o…