2 papers
cs.CV2026
RoLA: Rotary-Positioned Low-Rank Linear Attention for Efficient Diffusion Transformers
Zekun Zhang, Yixiang Cai, Yuxi Liu +9
Diffusion Transformers (DiTs) achieve strong video generation quality, but their dense spatiotemporal self-attention scales quadratically with sequence length and quickly becomes t…
cs.LG2025
Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
Zhoutong Wu, Yuan Zhang, Yiming Dong +4
Transformer models have driven breakthroughs across various language tasks by their strong capability to learn rich contextual representations. Scaling them to improve representati…