activity
20222024
most citedPushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction Tuning

11 citations · 51 across the 12 of their papers we have counts for

collaborators
Showing cs.CLShow all

12 papers · 1 filter

cs.CL2024

Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts

Nikolas Gritsch, Qizhen Zhang, Acyr Locatelli +2

Efficiency, specialization, and adaptability to new data distributions are qualities that are hard to combine in current Large Language Models. The Mixture of Experts (MoE) archite…

cs.CL2024

Multilingual Arbitrage: Optimizing Data Pools to Accelerate Multilingual Progress

Ayomide Odumakinde, Daniel D'souza, Pat Verga +2

The use of synthetic data has played a critical role in recent state-of-art breakthroughs. However, overly relying on a single oracle teacher model to generate data has been shown…

cs.CL20243 cited

To Code, or Not To Code? Exploring Impact of Code in Pre-training

Viraat Aryabumi, Yixuan Su, Raymond Ma +6

Including code in the pre-training data mixture, even for models not specifically designed for code, has become a common practice in LLMs pre-training. While there has been anecdot…

cs.CL202413 cited

Consent in Crisis: The Rapid Decline of the AI Data Commons

Shayne Longpre, Robert Mahari, Ariel Lee +46

General-purpose artificial intelligence (AI) systems are built on massive swathes of public web data, assembled into corpora such as C4, RefinedWeb, and Dolma. To our knowledge, we…

cs.CL2024

LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives

Luísa Shimabucoro, Sebastian Ruder, Julia Kreutzer +2

The widespread adoption of synthetic data raises new questions about how models generating the data can influence other large language models (LLMs) via distilled data. To start, o…

cs.CL20249 cited

Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model

Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong +14

Recent breakthroughs in large language models (LLMs) have centered around a handful of data-rich languages. What does it take to broaden access to breakthroughs beyond first-class…