activity
20242026
collaborators
Showing stat.MLShow all

6 papers · 1 filter

stat.ML2026

Fitted Occupancy-Ratio Evaluation without Bellman Completeness

Lars van der Laan, Nathan Kallus

Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate…

stat.ML2025

Soft Fitted Q-Iteration without Bellman Completeness: Occupancy Reweighting and Temperature Annealing

Lars van der Laan, Nathan Kallus

Fitted \(Q\)-iteration (FQI) is a standard regression-based method for optimal control in offline reinforcement learning, but its stability under function approximation often relie…

stat.ML2025

Fitted Q-Evaluation without Bellman Completeness via Occupancy Weighting

Lars van der Laan, Nathan Kallus

Fitted \(Q\)-evaluation (FQE) is a standard regression-based method for off-policy evaluation, but under distribution shift, value-function realizability alone does not ensure conv…

stat.ML2025

Bellman Calibration for -Learning in Offline Reinforcement Learning

Lars van der Laan, Nathan Kallus

Reliable long-horizon value prediction is difficult in offline reinforcement learning because fitted value methods combine bootstrapping, function approximation, and distribution s…

stat.ML2025

Semiparametric Double Reinforcement Learning with Applications to Long-Term Causal Inference

Lars van der Laan, David Hubbard, Allen Tran +2

Double reinforcement learning (DRL) provides efficient off-policy inference for policy values in nonparametric Markov decision processes (MDPs), but fully nonparametric estimators…

stat.ML2024

Self-Calibrating Conformal Prediction

Lars van der Laan, Ahmed M. Alaa

In machine learning, model calibration and predictive inference are essential for producing reliable predictions and quantifying uncertainty to support decision-making. Recognizing…