2 papers
cs.LG2026
TROLL: Trust Regions improve Reinforcement Learning for Large Language Models
Philipp Becker, Niklas Freymuth, Serge Thilges +2
Reinforcement Learning (RL) with PPO-like clip objectives has become the standard choice for reward-based fine-tuning of large language models (LLMs). Although recent work has expl…
cs.LG2025
Efficient Off-Policy Learning for High-Dimensional Action Spaces
Fabian Otto, Philipp Becker, Ngo Anh Vien +1
Existing off-policy reinforcement learning algorithms often rely on an explicit state-action-value function representation, which can be problematic in high-dimensional action spac…