16 citations · 36 across the 6 of their papers we have counts for
7 papers
MiniMax Sparse Attention
Xunhao Lai, Weiqi Xu, Yufeng Yang +14
Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointl…
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
Aili Chen, Aonian Li, Baichuan Zhou +215
We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
MiniMax, :, Aili Chen +125
We introduce MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. MiniMax-M1 is powered by a hybrid Mixture-of-Experts (MoE) architecture combin…
SmartTrim: Adaptive Tokens and Attention Pruning for Efficient Vision-Language Models
Zekun Wang, Jingchang Chen, Wangchunshu Zhou +7
Despite achieving remarkable performance on various vision-language tasks, Transformer-based Vision-Language Models (VLMs) suffer from redundancy in inputs and parameters, signific…
Learning to Tokenize for Generative Retrieval
Weiwei Sun, Lingyong Yan, Zheng Chen +7
Conventional document retrieval techniques are mainly based on the index-retrieve paradigm. It is challenging to optimize pipelines based on this paradigm in an end-to-end manner.…
HandAugment: A Simple Data Augmentation Method for Depth-Based 3D Hand Pose Estimation
Zhaohui Zhang, Shipeng Xie, Mingxiu Chen +1
Hand pose estimation from 3D depth images, has been explored widely using various kinds of techniques in the field of computer vision. Though, deep learning based method improve th…