3 papers
cs.CL2025
Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation
Israfel Salazar, Manuel Fernández Burda, Shayekh Bin Islam +42
The evaluation of vision-language models (VLMs) has mainly relied on English-language benchmarks, leaving significant gaps in both multilingual and multicultural coverage. While mu…
cs.CL2025
A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics
Luisa Shimabucoro, Ahmet Ustun, Marzieh Fadaee +1
In order for large language models to be useful across the globe, they are fine-tuned to follow instructions on multilingual data. Despite the ubiquity of such post-training, a cle…
cs.CL2024
LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives
LuÃsa Shimabucoro, Sebastian Ruder, Julia Kreutzer +2
The widespread adoption of synthetic data raises new questions about how models generating the data can influence other large language models (LLMs) via distilled data. To start, o…