4 papers · 1 filter
Mitigating Partial Observability in Sequential Decision Processes via the Lambda Discrepancy
Cameron Allen, Aaron Kirtland, Ruo Yu Tao +7
Reinforcement learning algorithms typically rely on the assumption that the environment dynamics and value function can be expressed in terms of a Markovian state representation. H…
An Optimal Tightness Bound for the Simulation Lemma
Sam Lobel, Ronald Parr
We present a bound for value-prediction error with respect to model misspecification that is tight, including constant factors. This is a direct improvement of the "simulation lemm…
Coarse-Grained Smoothness for RL in Metric Spaces
Omer Gottesman, Kavosh Asadi, Cameron Allen +3
Principled decision-making in continuous state--action spaces is impossible without some assumptions. A common approach is to assume Lipschitz continuity of the Q-function. We show…
Towards Amortized Ranking-Critical Training for Collaborative Filtering
Sam Lobel, Chunyuan Li, Jianfeng Gao +1
Collaborative filtering is widely used in modern recommender systems. Recent research shows that variational autoencoders (VAEs) yield state-of-the-art performance by integrating f…