11 citations · 18 across the 5 of their papers we have counts for
3 papers · 1 filter
Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback
Keertana Chidambaram, Sanath Kumar Krishnamurthy, Qiuling Xu +2
In recommendation systems, users interact with only a small fraction of a vast item catalog, producing feedback that is both sparse and noisy. This challenges post-training generat…
Robust Post-Training for Generative Recommenders: Why Exponential Reward-Weighted SFT Outperforms RLHF
Keertana Chidambaram, Sanath Kumar Krishnamurthy, Qiuling Xu +2
Aligning generative recommender systems to user preferences via post-training is critical for closing the gap between next-item prediction and actual recommendation quality. Existi…
Adaptive Exploration for Latent-State Bandits
Jikai Jin, Kenneth Hung, Sanath Kumar Krishnamurthy +2
We study bandits whose rewards depend on an unobserved Markov state that evolves independently of the learner's actions. The optimal arm can change even though the learner observes…