most citedMiMo-V2-Flash Technical Report

1 citations · 2 across the 6 of their papers we have counts for

collaborators

6 papers

cs.LG2026

Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation

Wenhui Chen, Zhifeng Li, Jie Zhou +5

A compressed student has two shapes that need not agree: the weight it deploys at inference and the weight family its training can reach. We show that a state-of-the-art weight-inh…

cs.AI2026

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud

Chenghua Wang, Daliang Xu, Dongqi Cai +23

Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Altho…

cs.CL2026

NanoSpec: Accelerating Speculative Decoding using Minimalist In-Context Vocabularies

Zhiyang Chen, Daliang Xu, Yinyuan Zhang +3

The massive vocabulary sizes of large language models, often exceeding 100k tokens, impose a computational bottleneck on the final linear projection layer during speculative decodi…

cs.LG2026

Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization

Jinghe Zhang, Daliang Xu, Chenghua Wang +5

Large language models (LLMs) are increasingly deployed on mobile devices, where Neural Processing Units (NPUs) necessitate fully static quantization for optimal inference efficienc…

cs.CL20261 cited

MiMo-V2-Flash Technical Report

Core Team, Bangjun Xiao, Bingquan Xia +123

We present MiMo-V2-Flash, a Mixture-of-Experts (MoE) model with 309B total parameters and 15B active parameters, designed for fast, strong reasoning and agentic capabilities. MiMo-…

cs.CL20241 cited

PhoneLM:an Efficient and Capable Small Language Model Family through Principled Pre-training

Rongjie Yi, Xiang Li, Weikai Xie +6

The interest in developing small language models (SLM) for on-device deployment is fast growing. However, the existing SLM design hardly considers the device hardware characteristi…