2 papers
cs.LG2026
Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching
Xiang Li, Nan Jiang
We present a novel theoretical framework, Q-MMR, for off-policy evaluation in finite-horizon MDPs. Q-MMR learns a set of scalar weights, one for each data point, such that the rewe…
math.OC2026
Tightening CVaR Approximations via Scenario-Wise Scaling for Chance-Constrained Programming
Rui Chen, Nan Jiang
Chance-constrained programs (CCPs) provide a powerful modeling framework for decision-making under uncertainty, but their nonconvex feasible regions make them computationally chall…