87 citations · 139 across the 10 of their papers we have counts for
Showing 2024 · cs.CLShow all
2 papers · 2 filters
cs.CL2024★ 1 cited
Diversify and Conquer: Diversity-Centric Data Selection with Iterative Refinement
Simon Yu, Liangyu Chen, Sara Ahmadian +1
Finetuning large language models on instruction data is crucial for enhancing pre-trained knowledge and improving instruction-following capabilities. As instruction datasets prolif…
cs.CL2024★ 1 cited
Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code
Taishi Nakamura, Mayank Mishra, Simone Tedeschi +42
Pretrained language models are an integral part of AI applications, but their high computational cost for training limits accessibility. Initiatives such as Bloom and StarCoder aim…