18 papers
Semiparametric inference on identification sets in choice modeling
Antoine Scheid, Jia Wan, Guy Aridor +2
In a discrete choice model, choice probabilities observed for a finite collection of choice sets may not identify a counterfactual choice probability under an unobserved choice set…
Fitted Occupancy-Ratio Evaluation without Bellman Completeness
Lars van der Laan, Nathan Kallus
Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate…
Mult-DPO: Multinomial Direct Preference Optimization for Recommender Systems
Yaochen Zhu, Harald Steck, James McInerney +4
Direct preference optimization (DPO) is a simple and effective alignment strategy for large language models (LLMs) based on pairwise preferences. In recommender systems, however, u…
The Value of Personalized Recommendations: Evidence from Netflix
Kevin Zielnicki, Guy Aridor, Aurélien Bibaut +3
Personalized recommendation systems shape much of user choice online, yet their targeted nature makes separating out the value of recommendation and the underlying goods challengin…
Stationary Reweighting Yields Local Convergence of Soft Fitted Q-Iteration
Lars van der Laan, Nathan Kallus
Fitted -iteration (FQI) and soft FQI are widely used value-based methods for offline reinforcement learning, but their standard stability guarantees often depend on Bellman comp…
Fitted Evaluation Without Bellman Completeness via Stationary Weighting
Lars van der Laan, Nathan Kallus
Fitted -evaluation (FQE) is a standard regression-based tool for off-policy evaluation, but existing stability guarantees often rely on Bellman completeness, a strong closure co…