5 papers
Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback
Keertana Chidambaram, Sanath Kumar Krishnamurthy, Qiuling Xu +2
In recommendation systems, users interact with only a small fraction of a vast item catalog, producing feedback that is both sparse and noisy. This challenges post-training generat…
Adaptive Exploration for Latent-State Bandits
Jikai Jin, Kenneth Hung, Sanath Kumar Krishnamurthy +2
We study bandits whose rewards depend on an unobserved Markov state that evolves independently of the learner's actions. The optimal arm can change even though the learner observes…
Robust Post-Training for Generative Recommenders: Why Exponential Reward-Weighted SFT Outperforms RLHF
Keertana Chidambaram, Sanath Kumar Krishnamurthy, Qiuling Xu +2
Aligning generative recommender systems to user preferences via post-training is critical for closing the gap between next-item prediction and actual recommendation quality. Existi…
Data-driven Error Estimation: Excess Risk Bounds without Class Complexity as Input
Sanath Kumar Krishnamurthy, Anna Lyubarskaja, Emma Brunskill +1
Constructing confidence intervals that are simultaneously valid across a class of estimates is central to tasks such as multiple mean estimation, generalization guarantees, and ada…
Selective Uncertainty Propagation in Offline RL
Sanath Kumar Krishnamurthy, Tanmay Gangwani, Sumeet Katariya +3
We consider the finite-horizon offline reinforcement learning (RL) setting, and are motivated by the challenge of learning the policy at any step h in dynamic programming (DP) algo…