activity
20232026
most citedLightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models

2 citations · 4 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL20251 cited

MiniMax-01: Scaling Foundation Models with Lightning Attention

MiniMax, Aonian Li, Bangwei Gong +87

We introduce MiniMax-01 series, including MiniMax-Text-01 and MiniMax-VL-01, which are comparable to top-tier models while offering superior capabilities in processing longer conte…

cs.CL2024

LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training

Xiaoye Qu, Daize Dong, Xuyang Hu +3

Recently, inspired by the concept of sparsity, Mixture-of-Experts (MoE) models have gained increasing popularity for scaling model size while keeping the number of activated parame…

cs.CL2024

Scaling Laws for Linear Complexity Language Models

Xuyang Shen, Dong Li, Ruitao Leng +3

The interest in linear complexity models for large language models is on the rise, although their scaling capacity remains uncertain. In this study, we present the scaling laws for…

cs.CL2024

Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention

Zhen Qin, Weigao Sun, Dong Li +3

We present Lightning Attention, the first linear attention implementation that maintains a constant training speed for various sequence lengths under fixed memory consumption. Due…

cs.CL2024

Unlocking the Secrets of Linear Complexity Sequence Model from A Unified Perspective

Zhen Qin, Xuyang Shen, Dong Li +4

We present the Linear Complexity Sequence Model (LCSM), a comprehensive solution that unites various sequence modeling techniques with linear complexity, including linear attention…

cs.CL2024

HGRN2: Gated Linear RNNs with State Expansion

Zhen Qin, Songlin Yang, Weixuan Sun +4

Hierarchically gated linear RNN (HGRN, \citealt{HGRN}) has demonstrated competitive training speed and performance in language modeling while offering efficient inference. However,…