708 citations · 972 across the 25 of their papers we have counts for
5 papers · 1 filter
ReMem: Mutual Information-Aware Fine-tuning of Pretrained Vision Transformers for Effective Knowledge Distillation
Chengyu Dong, Huan Gui, Noveen Sachdeva +6
Knowledge distillation from pretrained visual representation models offers an effective approach to improve small, task-specific production models. However, the effectiveness of su…
How to Train Data-Efficient LLMs
Noveen Sachdeva, Benjamin Coleman, Wang-Cheng Kang +6
The training of large language models (LLMs) is expensive. In this paper, we study data-efficient approaches for pre-training LLMs, i.e., techniques that aim to optimize the Pareto…
Online Matching: A Real-time Bandit System for Large-scale Recommendations
Xinyang Yi, Shao-Chuan Wang, Ruining He +6
The last decade has witnessed many successes of deep learning-based models for industry-scale recommender systems. These models are typically trained offline in a batch manner. Whi…
What Are Effective Labels for Augmented Data? Improving Calibration and Robustness with AutoLabel
Yao Qin, Xuezhi Wang, Balaji Lakshminarayanan +2
A wide breadth of research has devised data augmentation approaches that can improve both accuracy and generalization performance for neural networks. However, augmented data can e…
Improving Training Stability for Multitask Ranking Models in Recommender Systems
Jiaxi Tang, Yoel Drori, Daryl Chang +6
Recommender systems play an important role in many content platforms. While most recommendation research is dedicated to designing better models to improve user experience, we foun…