6 papers
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
Mix, MinHash, and Match: Cross-Source Agreement for Multilingual Pretraining Datasets
Sultan Alrashed, Francesco Orabona
Multilingual data from the web is essential for LLM pretraining. Yet, scraping it is expensive, and research groups repeatedly crawl the same content. For example, we found that ov…
SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
Sultan Alrashed, Chadi Helwe, Francesco Orabona
Although the community has tackled the acquisition of high-quality Arabic pretraining data, we still lack large-scale, multi-turn Arabic datasets that include reasoning and tool ca…
Beyond the Ideal: Analyzing the Inexact Muon Update
Egor Shulgin, Sultan AlRashed, Francesco Orabona +1
The Muon optimizer has rapidly emerged as a powerful, geometry-aware alternative to AdamW, demonstrating strong performance in large-scale training of neural networks. However, a c…
SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs
Sultan Alrashed
We present SmolTulu-1.7b-Instruct, referenced in this report as SmolTulu-DPO-1130, an instruction-tuned language model that adapts AllenAI's Tulu 3 post-training pipeline to enhanc…
Fineweb-Edu-Ar: Machine-translated Corpus to Support Arabic Small Language Models
Sultan Alrashed, Dmitrii Khizbullin, David R. Pugh
As large language models (LLMs) grow and develop, so do their data demands. This is especially true for multilingual LLMs, where the scarcity of high-quality and readily available…