2 papers
cs.LG2026
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits
Yuta Natsubori, Masataka Ushiku, Yuta Saito
Off-Policy Evaluation and Learning (OPE/L) in contextual bandits is rapidly gaining popularity in real systems because new policies can be evaluated and learned securely using only…
cs.LG2026
A More Accurate Algorithm Comparison through A/B Testing using Offline Evaluation Methods
Koki Konishi, Masataka Ushiku, Yuta Saito
A/B testing is the gold standard for selecting the better algorithm in online services. While offline evaluation has attracted attention as a safer alternative due to the high expe…