Tailored-LLaMA: Optimizing Few-Shot Learning in Pruned LLaMA Models with Task-Specific Prompts
arXiv:2410.19185 · doi:10.3233/FAIA392
Abstract
Large language models demonstrate impressive proficiency in language understanding and generation. Nonetheless, training these models from scratch, even the least complex billion-parameter variant demands significant computational resources rendering it economically impractical for many organizations. With large language models functioning as general-purpose task solvers, this paper investigates their task-specific fine-tuning. We employ task-specific datasets and prompts to fine-tune two pruned LLaMA models having 5 billion and 4 billion parameters. This process utilizes the pre-trained weights and focuses on a subset of weights using the LoRA method. One challenge in fine-tuning the LLaMA model is crafting a precise prompt tailored to the specific task. To address this, we propose a novel approach to fine-tune the LLaMA model under two primary constraints: task specificity and prompt effectiveness. Our approach, Tailored LLaMA initially employs structural pruning to reduce the model sizes from 7B to 5B and 4B parameters. Subsequently, it applies a carefully designed prompt specific to the task and utilizes the LoRA method to accelerate the fine-tuning process. Moreover, fine-tuning a model pruned by 50\% for less than one hour restores the mean accuracy of classification tasks to 95.68\% at a 20\% compression ratio and to 86.54\% at a 50\% compression ratio through few-shot learning with 50 shots. Our validation of Tailored LLaMA on these two pruned variants demonstrates that even when compressed to 50\%, the models maintain over 65\% of the baseline model accuracy in few-shot classification and generation tasks. These findings highlight the efficacy of our tailored approach in maintaining high performance with significantly reduced model sizes.
References in corpus (17)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- LLaMA: Open and Efficient Foundation Language Models
- LoRA: Low-Rank Adaptation of Large Language Models
- PaLM: Scaling Language Modeling with Pathways
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
- LaMDA: Language Models for Dialog Applications
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- DynaBERT: Dynamic BERT with Adaptive Width and Depth
- The Falcon Series of Open Language Models
- LLM-Pruner: On the Structural Pruning of Large Language Models
- SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
- A Simple and Effective Pruning Approach for Large Language Models
- Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods
- ZipLM: Inference-Aware Structured Pruning of Language Models
- Re-Reading Improves Reasoning in Large Language Models
- LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning
- Fluctuation-based Adaptive Structured Pruning for Large Language Models