167 citations · 567 across the 63 of their papers we have counts for
10 papers · 2 filters
Multi-Task Off-Policy Learning from Bandit Feedback
Joey Hong, Branislav Kveton, Sumeet Katariya +2
Many practical applications, such as recommender systems and learning to rank, involve solving multiple similar tasks. One example is learning of recommendation policies for users…
Differentially Private Adaptive Optimization with Delayed Preconditioners
Tian Li, Manzil Zaheer, Ken Ziyu Liu +3
Privacy noise may negate the benefits of using adaptive optimizers in differentially private model training. Prior works typically address this issue by using auxiliary information…
Learning to Navigate Wikipedia by Taking Random Walks
Manzil Zaheer, Kenneth Marino, Will Grathwohl +7
A fundamental ability of an intelligent web-based agent is seeking out and acquiring new information. Internet search engines reliably find the correct vicinity but the top results…
Generalization Properties of Retrieval-based Models
Soumya Basu, Ankit Singh Rawat, Manzil Zaheer
Many modern high-performing machine learning models such as GPT-3 primarily rely on scaling up models, e.g., transformer networks. Simultaneously, a parallel line of work aims to i…
A Fourier Approach to Mixture Learning
Mingda Qiao, Guru Guruganesh, Ankit Singh Rawat +2
We revisit the problem of learning mixtures of spherical Gaussians. Given samples from mixture , the goal is to estimate the means $…
Teacher Guided Training: An Efficient Framework for Knowledge Transfer
Manzil Zaheer, Ankit Singh Rawat, Seungyeon Kim +5
The remarkable performance gains realized by large pretrained models, e.g., GPT-3, hinge on the massive amounts of data they are exposed to during training. Analogously, distilling…