collaborators

6 papers

cs.CL2026

HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing

Yizhao Gao, Jianyu Wei, Qihao Zhang +11

This work introduces Hybrid Sparse Attention (HySparse), a new architecture that interleaves each full attention layer with several sparse attention layers. While conceptually simp…

cs.CL2026

MiMo-V2-Flash Technical Report

Core Team, Bangjun Xiao, Bingquan Xia +123

We present MiMo-V2-Flash, a Mixture-of-Experts (MoE) model with 309B total parameters and 15B active parameters, designed for fast, strong reasoning and agentic capabilities. MiMo-…

cs.CL2025

MiMo-Audio: Audio Language Models are Few-Shot Learners

Core Team, Dong Zhang, Gang Wang +97

Existing audio language models typically rely on task-specific fine-tuning to accomplish particular audio tasks. In contrast, humans are able to generalize to new audio tasks with…

cs.CL2025

Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers

Wenhan Ma, Hailin Zhang, Liang Zhao +4

Reinforcement learning (RL) has emerged as a crucial approach for enhancing the capabilities of large language models. However, in Mixture-of-Experts (MoE) models, the routing mech…

cs.CL2025

MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining

LLM-Core Xiaomi, :, Bingquan Xia +62

We present MiMo-7B, a large language model born for reasoning tasks, with optimization across both pre-training and post-training stages. During pre-training, we enhance the data p…

cs.CL2025

MiMo-VL Technical Report

Core Team, Zihao Yue, Zhenru Lin +71

We open-source MiMo-VL-7B-SFT and MiMo-VL-7B-RL, two powerful vision-language models delivering state-of-the-art performance in both general visual understanding and multimodal rea…