Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning
arXiv:2205.05638
Abstract
Few-shot in-context learning (ICL) enables pre-trained language models to perform a previously-unseen task without any gradient-based training by feeding a small number of training examples as part of the input. ICL incurs substantial computational, memory, and storage costs because it involves processing all of the training examples every time a prediction is made. Parameter-efficient fine-tuning (PEFT) (e.g. adapter modules, prompt tuning, sparse update methods, etc.) offers an alternative paradigm where a small set of parameters are trained to enable a model to perform the new task. In this paper, we rigorously compare few-shot ICL and PEFT and demonstrate that the latter offers better accuracy as well as dramatically lower computational costs. Along the way, we introduce a new PEFT method called (IA) that scales activations by learned vectors, attaining stronger performance while only introducing a relatively tiny amount of new parameters. We also propose a simple recipe based on the T0 model called T-Few that can be applied to new tasks without task-specific tuning or modifications. We validate the effectiveness of T-Few on completely unseen tasks by applying it to the RAFT benchmark, attaining super-human performance for the first time and outperforming the state-of-the-art by 6% absolute. All of the code used in our experiments is publicly available.
Cited by in corpus (21)
- Understanding Large-Language Model (LLM)-powered Human-Robot Interaction
- Advances of Machine Learning in Materials Science: Ideas and Techniques
- LLMParser: An Exploratory Study on Using Large Language Models for Log Parsing
- LLM4PLC: Harnessing Large Language Models for Verifiable Programming of PLCs in Industrial Control Systems
- Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
- Exploring the Roles of Large Language Models in Reshaping Transportation Systems: A Survey, Framework, and Roadmap
- Extracting Social Support and Social Isolation Information from Clinical Psychiatry Notes: Comparing a Rule-based NLP System and a Large Language Model
- DeTiME: Diffusion-Enhanced Topic Modeling using Encoder-decoder based LLM
- Opportunities and Challenges of Generative-AI in Finance
- A Few-Shot Approach to Dysarthric Speech Intelligibility Level Classification Using Transformers
- Soft Prompt Decoding for Multilingual Dense Retrieval
- CANAL -- Cyber Activity News Alerting Language Model: Empirical Approach vs. Expensive LLM
- HiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights
- FANAL -- Financial Activity News Alerting Language Modeling Framework
- Multi-Head Adapter Routing for Cross-Task Generalization
- FlanEC: Exploring Flan-T5 for Post-ASR Error Correction
- In a Few Words: Comparing Weak Supervision and LLMs for Short Query Intent Classification
- DeepClair: Utilizing Market Forecasts for Effective Portfolio Selection
- Leveraging Parameter Efficient Training Methods for Low Resource Text Classification: A Case Study in Marathi
- ModuLoRA: Finetuning 2-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers
- Transductive Learning for Textual Few-Shot Classification in API-based Embedding Models