Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning
Allen Nie, Anirudhan Badrinath, Nicholas Tomlin +5
Learning and skill mastery require extensive and deliberate practice. In many learning settings, producing high-quality pedagogical materials can require a high level of domain exp…
cs.AI2026
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
Stephane Hatgis-Kessell, Logan Mondal Bhamidipaty, Emma Brunskill
Human-designed reward functions for reinforcement learning (RL) agents are frequently misaligned with the humans' true, unobservable objectives, and thus act only as proxies. Optim…
cs.AI2025
Predicting Long Term Sequential Policy Value Using Softer Surrogates
Hyunji Nam, Allen Nie, Ge Gao +2
Off-policy policy evaluation (OPE) estimates the outcome of a new policy using historical data collected from a different policy. However, existing OPE methods cannot handle cases…