1 citations · 3 across the 5 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Attention Residuals
Kimi Team, Guangyu Chen, Yu Zhang +34
Residual connections with PreNorm are standard in modern LLMs, yet they accumulate all layer outputs with fixed unit weights. This uniform aggregation causes uncontrolled hidden-st…
cs.CL2025
Kimi Linear: An Expressive, Efficient Attention Architecture
Kimi Team, Yu Zhang, Zongyu Lin +57
We introduce Kimi Linear, a hybrid linear attention architecture that, for the first time, outperforms full attention under fair comparisons across various scenarios -- including s…
cs.CL2020
CLUE: A Chinese Language Understanding Evaluation Benchmark
Liang Xu, Hai Hu, Xuanwei Zhang +29
The advent of natural language understanding (NLU) benchmarks for English, such as GLUE and SuperGLUE allows new NLU models to be evaluated across a diverse set of tasks. These com…