3 papers
cs.LG2026
The Value of Mechanistic Priors in Sequential Decision Making
Itai Shufaro, Gal Benor, Shie Mannor
Hybrid mechanistic models, physical priors with learned residuals, promise to reduce the data required for good decisions, but have no computable criterion to test this. We charact…
cs.LG2024
On Bits and Bandits: Quantifying the Regret-Information Trade-off
Itai Shufaro, Nadav Merlis, Nir Weinberger +1
In many sequential decision problems, an agent performs a repeated task. He then suffers regret and obtains information that he may use in the following rounds. However, sometimes…
cs.LG2024
On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
Navdeep Kumar, Yashaswini Murthy, Itai Shufaro +3
We present the first finite time global convergence analysis of policy gradient in the context of infinite horizon average reward Markov decision processes (MDPs). Specifically, we…