5 papers
BACH: A Bayesian Admixture of Contrastive Heads for Multi-Interest Two-Tower Retrieval
Quoc Phong Nguyen, Paul Albert, Long Vuong +2
Two-tower retrievers compress each user into a single embedding, limiting their ability to serve diverse interests. Multi-interest models give each user several heads scored by a m…
SineLoRA: Sine-Activated Delta Compression
Cameron Gordon, Yiping Ji, Hemanth Saratchandran +2
Resource-constrained weight deployment is a task of immense practical importance. Recently, there has been interest in the specific task of \textit{Delta Compression}, where partie…
Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product
Paul Albert, Frederic Z. Zhang, Hemanth Saratchandran +2
Parameter-efficient fine-tuning (PEFT) has become a standard approach for adapting large pre-trained models. Amongst PEFT methods, low-rank adaptation (LoRA) has achieved notable s…
RandLoRA: Full-rank parameter-efficient fine-tuning of large models
Paul Albert, Frederic Z. Zhang, Hemanth Saratchandran +3
Low-Rank Adaptation (LoRA) and its variants have shown impressive results in reducing the number of trainable parameters and memory requirements of large transformer networks while…
Knowledge Composition using Task Vectors with Learned Anisotropic Scaling
Frederic Z. Zhang, Paul Albert, Cristian Rodriguez-Opazo +2
Pre-trained models produce strong generic representations that can be adapted via fine-tuning. The learned weight difference relative to the pre-trained model, known as a task vect…