2 papers
stat.ML2024
Decision-Centered Abstractions via Orthogonal Estimation of Difference-of-Q Functions
Defu Cao, Angela Zhou
Offline reinforcement learning enables evaluation and optimization of sequential decisions from historical data, when it is not possible to deploy new policies online due to safety…
stat.ML2024
Reward-Relevance-Filtered Linear Offline Reinforcement Learning
Angela Zhou
This paper studies offline reinforcement learning with linear function approximation in a setting with decision-theoretic, but not estimation sparsity. The structural restrictions…