activity
20172025
most citedScalable and Efficient MoE Training for Multitask Multilingual Models

33 citations · 63 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CL2025

Command A: An Enterprise-Ready Large Language Model

Team Cohere, :, Aakanksha +227

In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command A is an agent-optimised…

cs.LG2024

u-P: The Unit-Scaled Maximal Update Parametrization

Charlie Blake, Constantin Eichenberg, Josef Dean +7

The Maximal Update Parametrization (P) aims to make the optimal hyperparameters (HPs) of a model independent of its size, allowing them to be swept using a cheap proxy model rat…

cs.CV2023★ 6 cited

MultiFusion: Fusing Pre-Trained Models for Multi-Lingual, Multi-Modal Image Generation

Marco Bellagente, Manuel Brack, Hannah Teufel +11

The recent popularity of text-to-image diffusion models (DM) can largely be attributed to the intuitive interface they provide to users. The intended generation can be expressed in…

cs.CV2022★ 7 cited

M-VADER: A Model for Diffusion with Multimodal Context

Samuel Weinbach, Marco Bellagente, Constantin Eichenberg +7

We introduce M-VADER: a diffusion model (DM) for image generation where the output can be specified using arbitrary combinations of images and text. We show how M-VADER enables the…

cs.CL2022★ 4 cited

Building a great multi-lingual teacher with sparsely-gated mixture of experts for speech recognition

Kenichi Kumatani, Robert Gmyr, Felipe Cruz Salinas +5

The sparsely-gated Mixture of Experts (MoE) can magnify a network capacity with a little computational complexity. In this work, we investigate how multi-lingual Automatic Speech R…

cs.CL2021★ 33 cited

Scalable and Efficient MoE Training for Multitask Multilingual Models

Young Jin Kim, Ammar Ahmad Awan, Alexandre Muzio +6

The Mixture of Experts (MoE) models are an emerging class of sparsely activated deep learning models that have sublinear compute costs with respect to their parameters. In contrast…