Revisiting Instruction Fine-tuned Model Evaluation to Guide Industrial Applications
arXiv:2310.14103 · doi:10.18653/v1/2023.emnlp-main.559
Abstract
Instruction Fine-Tuning (IFT) is a powerful paradigm that strengthens the zero-shot capabilities of Large Language Models (LLMs), but in doing so induces new evaluation metric requirements. We show LLM-based metrics to be well adapted to these requirements, and leverage them to conduct an investigation of task-specialization strategies, quantifying the trade-offs that emerge in practical industrial settings. Our findings offer practitioners actionable insights for real-world IFT model deployment.
Short paper accepted at EMNLP 2023
References in corpus (15)
- Training language models to follow instructions with human feedback
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
- Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
- OpenAssistant Conversations -- Democratizing Large Language Model Alignment
- LIMA: Less Is More for Alignment
- Large Language Models Are State-of-the-Art Evaluators of Translation Quality
- WizardLM: Empowering large pre-trained language models to follow complex instructions
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
- What are the best systems? New perspectives on NLP Benchmarking
- Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
- Automatic Text Evaluation through the Lens of Wasserstein Barycenters