2 papers
cs.CL2026
PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning
Hang Zhang, Warren J. Gross
Not all training samples contribute equally to large language model fine-tuning. Selecting informative training samples can reduce the computational cost while preserving downstrea…
cs.CL2024
Automatic Pruning of Fine-tuning Datasets for Transformer-based Language Models
Mohammadreza Tayaranian, Seyyed Hasan Mozafari, Brett H. Meyer +2
Transformer-based language models have shown state-of-the-art performance on a variety of natural language understanding tasks. To achieve this performance, these models are first…