3 papers
cs.LG2026
Learning Correlated Reward Models: Statistical Barriers and Opportunities
Yeshwanth Cherapanamjeri, Constantinos Daskalakis, Gabriele Farina +1
Random Utility Models (RUMs) are a classical framework for modeling user preferences and play a key role in reward modeling for Reinforcement Learning from Human Feedback (RLHF). H…
cs.LG2026
Reevaluating Policy Gradient Methods for Imperfect-Information Games
Max Rudolph, Nathan Lichtle, Sobhan Mohammadpour +6
In the past decade, motivated by the putative failure of naive self-play deep reinforcement learning (DRL) in adversarial imperfect-information games, researchers have developed nu…
stat.ME2025
Arc travel time and path choice model estimation subsumed
Sobhan Mohammadpour, Emma Frejinger
We address the problem of simultaneously estimating arc travel times in a network \emph{and} parameters of route choice models for strategic and tactical network planning purposes.…