12 papers · 1 filter
The Art of Asking: Multilingual Prompt Optimization for Synthetic Data
David Mora, Viraat Aryabumi, Wei-Yin Ko +3
Synthetic data has become a cornerstone for scaling large language models, yet its multilingual use remains bottlenecked by translation-based prompts. This strategy inherits Englis…
Verification Limits Code LLM Training
Srishti Gureja, Elena Tommasone, Jingyi He +3
Large language models for code generation increasingly rely on synthetic data, where both problem solutions and verification tests are generated by models. While this enables scala…
When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs
Ammar Khairi, Daniel D'souza, Ye Shen +2
Recent advancements in large language models (LLMs) have shifted focus toward scaling inference-time compute, improving performance without retraining the model. A common approach…
Language Models can perform Single-Utterance Self-Correction of Perturbed Reasoning
Sam Silver, Jimin Sun, Ivan Zhang +2
Large Language Models (LLMs) have demonstrated impressive mathematical reasoning capabilities, yet their performance remains brittle to minor variations in problem description and…
The Multilingual Divide and Its Impact on Global AI Safety
Aidan Peppin, Julia Kreutzer, Alice Schoenauer Sebag +13
Despite advances in large language model capabilities in recent years, a large gap remains in their capabilities and safety performance for many languages beyond a relatively small…
M-RewardBench: Evaluating Reward Models in Multilingual Settings
Srishti Gureja, Lester James V. Miranda, Shayekh Bin Islam +7
Reward models (RMs) have driven the state-of-the-art performance of LLMs today by enabling the integration of human feedback into the language modeling process. However, RMs are pr…