3 papers
cs.CL2026
Erase-then-Delta Attention: Decoupling Erase and Write Addresses in Delta-Rule Linear Attention
Xiao Li, Chengruidong Zhang, Hao Luo +15
Delta-rule linear attention improves recurrent memory updates by correcting what is already stored at the current write address before writing new content. However, the active corr…
cs.CV2026
VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inference
Anmin Liu, Ruixuan Yang, Huiqiang Jiang +5
Long-context video understanding and generation pose a significant computational challenge for Transformer-based video models due to the quadratic complexity of self-attention. Whi…
cs.CL2025
Qwen2.5-1M Technical Report
An Yang, Bowen Yu, Chengyuan Li +25
We introduce Qwen2.5-1M, a series of models that extend the context length to 1 million tokens. Compared to the previous 128K version, the Qwen2.5-1M series have significantly enha…