Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Switch Attention: Towards Dynamic and Fine-grained Hybrid Transformers
Yusheng Zhao, Hourun Li, Bohan Wu +5
The attention mechanism has been the core component in modern transformer architectures. However, the computation of standard full attention scales quadratically with the sequence…
cs.CL2025
ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language Models
Kangjie Zheng, Junwei Yang, Siyue Liang +5
Masked Language Models (MLMs) have achieved remarkable success in many self-supervised representation learning tasks. MLMs are trained by randomly masking portions of the input seq…