activity
20242026
collaborators

11 papers

cs.LG2026

Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning

Lars van der Laan, Nathan Kallus

Marginalized importance weighting evaluates a target policy by reweighting offline state-action samples with its discounted occupancy ratio, characterized by an adjoint Bellman equ…

stat.ML2026

Fitted Occupancy-Ratio Evaluation without Bellman Completeness

Lars van der Laan, Nathan Kallus

Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate…

cs.LG2026

Reward Transfer from Inverse Reinforcement Learning: A Coupled Minimax Approach

Guang-Yuan Hao, Lars van der Laan, Aurélien Bibaut +1

We study the transfer of rewards learned using inverse reinforcement learning from expert demonstrations in one environment to reinforcement learning in a new, different environmen…

cs.LG2025

Efficient Inference for Inverse Reinforcement Learning and Dynamic Discrete Choice Models

Lars van der Laan, Aurélien Bibaut, Nathan Kallus

In many sequential decision-making problems, researchers observe actions but not the rewards that drive behavior, yet still wish to evaluate and compare counterfactual policies. In…

stat.ML2025

Soft Fitted Q-Iteration without Bellman Completeness: Occupancy Reweighting and Temperature Annealing

Lars van der Laan, Nathan Kallus

Fitted \(Q\)-iteration (FQI) is a standard regression-based method for optimal control in offline reinforcement learning, but its stability under function approximation often relie…

stat.ML2025

Fitted Q-Evaluation without Bellman Completeness via Occupancy Weighting

Lars van der Laan, Nathan Kallus

Fitted \(Q\)-evaluation (FQE) is a standard regression-based method for off-policy evaluation, but under distribution shift, value-function realizability alone does not ensure conv…