activity
20242026
most citedMiMo-V2-Flash Technical Report

1 citations · 1 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing

Yizhao Gao, Jianyu Wei, Qihao Zhang +11

This work introduces Hybrid Sparse Attention (HySparse), a new architecture that interleaves each full attention layer with several sparse attention layers. While conceptually simp…

cs.CL20261 cited

MiMo-V2-Flash Technical Report

Core Team, Bangjun Xiao, Bingquan Xia +123

We present MiMo-V2-Flash, a Mixture-of-Experts (MoE) model with 309B total parameters and 15B active parameters, designed for fast, strong reasoning and agentic capabilities. MiMo-…

cs.CL2025

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

MiniMax, :, Aili Chen +125

We introduce MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. MiniMax-M1 is powered by a hybrid Mixture-of-Experts (MoE) architecture combin…

cs.CL2025

SwiLTra-Bench: The Swiss Legal Translation Benchmark

Joel Niklaus, Jakob Merane, Luka Nenadic +16

In Switzerland legal translation is uniquely important due to the country's four official languages and requirements for multilingual legal documentation. However, this process tra…

cs.CL2025

MiniMax-01: Scaling Foundation Models with Lightning Attention

MiniMax, Aonian Li, Bangwei Gong +87

We introduce MiniMax-01 series, including MiniMax-Text-01 and MiniMax-VL-01, which are comparable to top-tier models while offering superior capabilities in processing longer conte…

cs.CL2024

Predictable Emergent Abilities of LLMs: Proxy Tasks Are All You Need

Bo-Wen Zhang, Yan Yan, Boxiang Yang +2

While scaling laws optimize training configurations for large language models (LLMs) through experiments on smaller or early-stage models, they fail to predict emergent abilities d…