1 paper
Yan Xie, Tiansheng Wen, Tangda Huang +4
Scaling Transformers to ultra-long contexts is bottlenecked by the O(n2d) cost of self-attention. Existing methods reduce this cost along the sequence axis through local window…