30 citations · 32 across the 4 of their papers we have counts for
5 papers
The Perfect Blend: Redefining RLHF with Mixture of Judges
Tengyu Xu, Eryk Helenowski, Karthik Abinav Sankararaman +17
Reinforcement learning from human feedback (RLHF) has become the leading approach for fine-tuning large language models (LLM). However, RLHF has limitations in multi-task learning…
AD-DROP: Attribution-Driven Dropout for Robust Language Model Fine-Tuning
Tao Yang, Jinghao Deng, Xiaojun Quan +2
Fine-tuning large pre-trained language models on downstream tasks is apt to suffer from overfitting when limited training data is available. While dropout proves to be an effective…
Detection, Disambiguation, Re-ranking: Autoregressive Entity Linking as a Multi-Task Problem
Khalil Mrini, Shaoliang Nie, Jiatao Gu +3
We propose an autoregressive entity linking model, that is trained with two auxiliary tasks, and learns to re-rank generated samples at inference time. Our proposed novelties addre…
MSD: Saliency-aware Knowledge Distillation for Multimodal Understanding
Woojeong Jin, Maziar Sanjabi, Shaoliang Nie +3
To reduce a model size but retain performance, we often rely on knowledge distillation (KD) which transfers knowledge from a large "teacher" model to a smaller "student" model. How…
High Resolution Face Completion with Multiple Controllable Attributes via Fully End-to-End Progressive Generative Adversarial Networks
Zeyuan Chen, Shaoliang Nie, Tianfu Wu +1
We present a deep learning approach for high resolution face completion with multiple controllable attributes (e.g., male and smiling) under arbitrary masks. Face completion entail…