2 papers
cs.LG2026
Attention Needs to Focus: A Unified Perspective on Attention Allocation
Zichuan Fu, Wentao Song, Guojing Li +6
The Transformer architecture, a cornerstone of modern Large Language Models (LLMs), has achieved extraordinary success in sequence modeling, primarily due to its attention mechanis…
cs.CL2025
Sliding Window Attention Training for Efficient Large Language Models
Zichuan Fu, Wentao Song, Yejing Wang +7
Recent advances in transformer-based Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their quadratic computational complexity…