2 papers
cs.LG2026
Representation-Driven Reinforcement Learning
Ofir Nabati, Guy Tennenholtz, Shie Mannor
We present a representation-driven framework for reinforcement learning. By representing policies as estimates of their expected values, we leverage techniques from contextual band…
cs.LG2025
Policy Gradient with Tree Expansion
Gal Dalal, Assaf Hallak, Gugan Thoppe +2
Policy gradient methods are notorious for having a large variance and high sample complexity. To mitigate this, we introduce SoftTreeMax -- a generalization of softmax that employs…