3 papers
cs.SE2025
Verification Limits Code LLM Training
Srishti Gureja, Elena Tommasone, Jingyi He +3
Large language models for code generation increasingly rely on synthetic data, where both problem solutions and verification tests are generated by models. While this enables scala…
cs.AI2025
The Multilingual Divide and Its Impact on Global AI Safety
Aidan Peppin, Julia Kreutzer, Alice Schoenauer Sebag +13
Despite advances in large language model capabilities in recent years, a large gap remains in their capabilities and safety performance for many languages beyond a relatively small…
cs.CL2025
Aya Vision: Advancing the Frontier of Multilingual Multimodality
Saurabh Dash, Yiyang Nan, John Dang +22
Building multimodal language models is fundamentally challenging: it requires aligning vision and language modalities, curating high-quality instruction data, and avoiding the degr…