2 papers
cs.LG2026
Power-based Partial Attention: Bridging Linear-Complexity and Full Attention
Yufeng Huang
It is widely accepted from transformer research that "attention is all we need", but the amount of attention required has never been systematically quantified. Is quadratic $O(L^2)…
cs.LG2026
Superlinear Multi-Step Attention
Yufeng Huang
In this paper, we propose \textbf{Superlinear attention}, a fully trainable multi-step attention architecture that achieves subquadratic complexity for long sequences while preserv…