most citedPushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction Tuning

11 citations · 38 across the 9 of their papers we have counts for

collaborators

9 papers

cs.CL20236 cited

The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI

Shayne Longpre, Robert Mahari, Anthony Chen +14

The race to train language models on vast, diverse, and inconsistently documented datasets has raised pressing concerns about the legal and ethical risks for practitioners. To reme…

cs.CL20234 cited

Which Prompts Make The Difference? Data Prioritization For Efficient Human LLM Evaluation

Meriem Boubdir, Edward Kim, Beyza Ermis +2

Human evaluation is increasingly critical for assessing large language models, capturing linguistic nuances, and reflecting user preferences more accurately than traditional automa…

cs.AI20231 cited

Goodtriever: Adaptive Toxicity Mitigation with Retrieval-augmented Models

Luiza Pozzobon, Beyza Ermis, Patrick Lewis +1

Considerable effort has been dedicated to mitigating toxicity, but existing methods often require drastic modifications to model parameters or the use of computationally intensive…

cs.SE20231 cited

The Grand Illusion: The Myth of Software Portability and Implications for ML Progress

Fraser Mince, Dzung Dinh, Jonas Kgomo +2

Pushing the boundaries of machine learning often requires exploring different hardware and software combinations. However, the freedom to experiment across different tooling stacks…

cs.CL202311 cited

Pushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction Tuning

Ted Zadouri, Ahmet Üstün, Arash Ahmadian +3

The Mixture of Experts (MoE) is a widely known neural architecture where an ensemble of specialized sub-models optimizes overall performance with a constant computational cost. How…

cs.CL20237 cited

When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Max Marion, Ahmet Üstün, Luiza Pozzobon +3

Large volumes of text data have contributed significantly to the development of large language models (LLMs) in recent years. This data is typically acquired by scraping the intern…