13 citations · 31 across the 8 of their papers we have counts for
11 papers
RewardRank: Optimizing True Learning-to-Rank Utility
Gaurav Bhatt, Kiran Koshy Thekumparampil, Tanmay Gangwani +2
Traditional ranking systems optimize offline proxy objectives that rely on oversimplified assumptions about user behavior, often neglecting factors such as position bias and item d…
Knowledge Distillation with Training Wheels
Guanlin Liu, Anand Ramachandran, Tanmay Gangwani +2
Knowledge distillation is used, in generative language modeling, to train a smaller student model using the help of a larger teacher model, resulting in improved capabilities for t…
Imitation Learning from Observations under Transition Model Disparity
Tanmay Gangwani, Yuan Zhou, Jian Peng
Learning to perform tasks by leveraging a dataset of expert observations, also known as imitation learning from observations (ILO), is an important paradigm for learning skills wit…
Harnessing Distribution Ratio Estimators for Learning Agents with Quality and Diversity
Tanmay Gangwani, Jian Peng, Yuan Zhou
Quality-Diversity (QD) is a concept from Neuroevolution with some intriguing applications to Reinforcement Learning. It facilitates learning a population of agents where each membe…
Learning Guidance Rewards with Trajectory-space Smoothing
Tanmay Gangwani, Yuan Zhou, Jian Peng
Long-term temporal credit assignment is an important challenge in deep reinforcement learning (RL). It refers to the ability of the agent to attribute actions to consequences that…
Mutual Information Based Knowledge Transfer Under State-Action Dimension Mismatch
Michael Wan, Tanmay Gangwani, Jian Peng
Deep reinforcement learning (RL) algorithms have achieved great success on a wide variety of sequential decision-making tasks. However, many of these algorithms suffer from high sa…