6 papers · 1 filter
Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier
John Dang, Shivalika Singh, Daniel D'souza +42
We introduce the Aya Expanse model family, a new generation of 8B and 32B parameter multilingual language models, aiming to address the critical challenge of developing highly perf…
Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning
Aakanksha, Arash Ahmadian, Seraphina Goldfarb-Tarrant +3
Large Language Models (LLMs) have been adopted and deployed worldwide for a broad variety of applications. However, ensuring their safe use remains a significant challenge. Prefere…
Multilingual Arbitrage: Optimizing Data Pools to Accelerate Multilingual Progress
Ayomide Odumakinde, Daniel D'souza, Pat Verga +2
The use of synthetic data has played a critical role in recent state-of-art breakthroughs. However, overly relying on a single oracle teacher model to generate data has been shown…
LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives
LuÃsa Shimabucoro, Sebastian Ruder, Julia Kreutzer +2
The widespread adoption of synthetic data raises new questions about how models generating the data can influence other large language models (LLMs) via distilled data. To start, o…
The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
Aakanksha, Arash Ahmadian, Beyza Ermis +4
A key concern with the concept of "alignment" is the implicit question of "alignment to what?". AI systems are increasingly used across the world, yet safety alignment is often foc…
From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models
Luiza Pozzobon, Patrick Lewis, Sara Hooker +1
To date, toxicity mitigation in language models has almost entirely been focused on single-language settings. As language models embrace multilingual capabilities, it's crucial our…