3 papers
cs.IR2026
Credit-assigned Policy Gradient for Early Stage Retrieval in Two-stage Ranking
Haruka Kiyohara, Mihaela Curmei, Ariel Evnine +7
Large-scale search, recommendation, and retrieval-augmented generation (RAG) systems typically employ a two-stage architecture: an early-stage ranker (ESR) generates a candidate se…
stat.ML2025
SNPL: Simultaneous Policy Learning and Evaluation for Safe Multi-Objective Policy Improvement
Brian Cho, Ana-Roxana Pop, Ariel Evnine +1
To design effective digital interventions, experimenters face the challenge of learning decision policies that balance multiple objectives using offline data. Often, they aim to de…
cs.LG2024
CSPI-MT: Calibrated Safe Policy Improvement with Multiple Testing for Threshold Policies
Brian M Cho, Ana-Roxana Pop, Kyra Gan +4
When modifying existing policies in high-risk settings, it is often necessary to ensure with high certainty that the newly proposed policy improves upon a baseline, such as the sta…