1 paper · 1 filter
Paramita Mirza, Lucas Weber, Fabian Küch
Recent work shows that post-training datasets for LLMs can be substantially downsampled without noticeably deteriorating performance. However, data selection often incurs high comp…