1 paper
Paramita Mirza, Lucas Weber, Fabian Küch
Recent work shows that post-training datasets for LLMs can be substantially downsampled without noticeably deteriorating performance. However, data selection often incurs high comp…