1 citations · 1 across the 4 of their papers we have counts for
4 papers · 1 filter
Mix, MinHash, and Match: Cross-Source Agreement for Multilingual Pretraining Datasets
Sultan Alrashed, Francesco Orabona
Multilingual data from the web is essential for LLM pretraining. Yet, scraping it is expensive, and research groups repeatedly crawl the same content. For example, we found that ov…
SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
Sultan Alrashed, Chadi Helwe, Francesco Orabona
Although the community has tackled the acquisition of high-quality Arabic pretraining data, we still lack large-scale, multi-turn Arabic datasets that include reasoning and tool ca…
SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs
Sultan Alrashed
We present SmolTulu-1.7b-Instruct, referenced in this report as SmolTulu-DPO-1130, an instruction-tuned language model that adapts AllenAI's Tulu 3 post-training pipeline to enhanc…
Fineweb-Edu-Ar: Machine-translated Corpus to Support Arabic Small Language Models
Sultan Alrashed, Dmitrii Khizbullin, David R. Pugh
As large language models (LLMs) grow and develop, so do their data demands. This is especially true for multilingual LLMs, where the scarcity of high-quality and readily available…