2 papers
cs.LG2026
Quantile of Means: A Bonus-Free Ensemble Method for Minimax Optimal Reinforcement Learning
Asaf Cassel, Aviv Rosenberg
Optimal Reinforcement Learning (RL) algorithms typically rely on carefully constructed count-based uncertainty estimates to drive exploration. Although theoretically sound, such es…
cs.LG2024
Multi-turn Reinforcement Learning from Preference Human Feedback
Lior Shani, Aviv Rosenberg, Asaf Cassel +10
Reinforcement Learning from Human Feedback (RLHF) has become the standard approach for aligning Large Language Models (LLMs) with human preferences, allowing LLMs to demonstrate re…