4 papers
RewardRank: Optimizing True Learning-to-Rank Utility
Gaurav Bhatt, Kiran Koshy Thekumparampil, Tanmay Gangwani +2
Traditional ranking systems optimize offline proxy objectives that rely on oversimplified assumptions about user behavior, often neglecting factors such as position bias and item d…
Knowledge Distillation with Training Wheels
Guanlin Liu, Anand Ramachandran, Tanmay Gangwani +2
Knowledge distillation is used, in generative language modeling, to train a smaller student model using the help of a larger teacher model, resulting in improved capabilities for t…
Selective Uncertainty Propagation in Offline RL
Sanath Kumar Krishnamurthy, Tanmay Gangwani, Sumeet Katariya +3
We consider the finite-horizon offline reinforcement learning (RL) setting, and are motivated by the challenge of learning the policy at any step h in dynamic programming (DP) algo…
Multi-Objective Optimization via Wasserstein-Fisher-Rao Gradient Flow
Yinuo Ren, Tesi Xiao, Tanmay Gangwani +4
Multi-objective optimization (MOO) aims to optimize multiple, possibly conflicting objectives with widespread applications. We introduce a novel interacting particle method for MOO…