3 papers
cs.LG2026
A Simple and Effective Reinforcement Learning Method for Text-to-Image Diffusion Fine-tuning
Shashank Gupta, Chaitanya Ahuja, Tsung-Yu Lin +4
Reinforcement learning (RL)-based fine-tuning has emerged as a powerful approach for aligning diffusion models with black-box objectives. Proximal policy optimization (PPO) is a po…
cs.IR2026
Towards Two-Stage Counterfactual Learning to Rank
Shashank Gupta, Yiming Liao, Maarten de Rijke
Counterfactual learning to rank (CLTR) aims to learn a ranking policy from user interactions while correcting for the inherent biases in interaction data, such as position bias. Ex…
cs.LG2024
A Simpler Alternative to Variational Regularized Counterfactual Risk Minimization
Hua Chang Bakker, Shashank Gupta, Harrie Oosterhuis
Variance regularized counterfactual risk minimization (VRCRM) has been proposed as an alternative off-policy learning (OPL) method. VRCRM method uses a lower-bound on the -diver…