activity
20222026
most citedMiMo-Audio: Audio Language Models are Few-Shot Learners

2 citations · 5 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

15 papers · 1 filter

cs.CL2026

Learning to Draft: Adaptive Speculative Decoding with Reinforcement Learning

Jiebin Zhang, Zhenghan Yu, Liang Wang +8

Speculative decoding accelerates large language model (LLM) inference by using a small draft model to generate candidate tokens for a larger target model to verify. The efficacy of…

cs.CL2026

PaperBanana: Automating Academic Illustration for AI Scientists

Dawei Zhu, Rui Meng, Yale Song +4

Despite rapid advances in autonomous AI scientists powered by language models, generating publication-ready illustrations remains a labor-intensive bottleneck in the research workf…

cs.CL20252 cited

MiMo-Audio: Audio Language Models are Few-Shot Learners

Core Team, Dong Zhang, Gang Wang +97

Existing audio language models typically rely on task-specific fine-tuning to accomplish particular audio tasks. In contrast, humans are able to generalize to new audio tasks with…

cs.CL2025

Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling

Junfeng Ran, Guangxiang Zhao, Yuhan Wu +7

The Mixture-of-Experts (MoE) models have gained significant attention in deep learning due to their dynamic resource allocation and superior performance across diverse tasks. Howev…

cs.CL2025

Hierarchical Memory Organization for Wikipedia Generation

Eugene J. Yu, Dawei Zhu, Yifan Song +6

Generating Wikipedia articles autonomously is a challenging task requiring the integration of accurate, comprehensive, and well-structured information from diverse sources. This pa…

cs.CL20251 cited

MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining

LLM-Core Xiaomi, :, Bingquan Xia +62

We present MiMo-7B, a large language model born for reasoning tasks, with optimization across both pre-training and post-training stages. During pre-training, we enhance the data p…