activity
20182026
most citedPushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction Tuning

11 citations · 36 across the 15 of their papers we have counts for

collaborators

19 papers

cs.CL2026

Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning

Kevin Du, Alexander Hoyle, Laura Ruis +1

Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges…

stat.ML2026

Efficient Benchmarking Is Just Feature Selection and Multiple Regression

Sam Bowyer, Acyr Locatelli, Kris Cao

Efficient benchmarking techniques aim to lower the computational cost of evaluating LLMs by predicting full benchmark scores using only a subset of a benchmark's questions. By refr…

cs.CL2026

Tiny Aya: Bridging Scale and Multilingual Depth

Alejandro R. Salamanca, Diana Abagyan, Daniel D'souza +23

Tiny Aya redefines what a small multilingual language model can achieve. Trained on 70 languages and refined through region-aware posttraining, it delivers state-of-the-art in tran…

cs.CL2025

One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers

Diana Abagyan, Alejandro R. Salamanca, Andres Felipe Cruz-Salinas +6

Pretraining massively multilingual Large Language Models (LLMs) for many languages at once is challenging due to limited model capacity, scarce high-quality data, and compute const…

cs.CL2025

Aya Vision: Advancing the Frontier of Multilingual Multimodality

Saurabh Dash, Yiyang Nan, John Dang +22

Building multimodal language models is fundamentally challenging: it requires aligning vision and language modalities, curating high-quality instruction data, and avoiding the degr…

cs.CL2025

Command A: An Enterprise-Ready Large Language Model

Team Cohere, :, Aakanksha +227

In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command A is an agent-optimised…