RoFormer: Enhanced Transformer with Rotary Position Embedding
arXiv:2104.09864
Abstract
Position encoding recently has shown effective in the transformer architecture. It enables valuable supervision for dependency modeling between elements at different positions of the sequence. In this paper, we first investigate various methods to integrate positional information into the learning process of transformer-based language models. Then, we propose a novel method named Rotary Position Embedding(RoPE) to effectively leverage the positional information. Specifically, the proposed RoPE encodes the absolute position with a rotation matrix and meanwhile incorporates the explicit relative position dependency in self-attention formulation. Notably, RoPE enables valuable properties, including the flexibility of sequence length, decaying inter-token dependency with increasing relative distances, and the capability of equipping the linear self-attention with relative position encoding. Finally, we evaluate the enhanced transformer with rotary position embedding, also called RoFormer, on various long text classification benchmark datasets. Our experiments show that it consistently overcomes its alternatives. Furthermore, we provide a theoretical analysis to explain some experimental results. RoFormer is already integrated into Huggingface: \url{https://huggingface.co/docs/transformers/model_doc/roformer}.
fixed some typos
References in corpus (8)
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
- fairseq: A Fast, Extensible Toolkit for Sequence Modeling
- Rethinking Attention with Performers
- Are Transformers universal approximators of sequence-to-sequence functions?
- CAIL2019-SCM: A Dataset of Similar Case Matching in Legal Domain
- Relative Positional Encoding for Transformers with Linear Complexity
- Translational Equivariance in Kernelizable Attention
Cited by in corpus (16)
- RiNALMo: General-Purpose RNA Language Models Can Generalize Well on Structure Prediction Tasks
- SemiCurv: Semi-Supervised Curvilinear Structure Segmentation
- A Large Encoder-Decoder Family of Foundation Models For Chemical Language
- Astroconformer: The Prospects of Analyzing Stellar Light Curves with Transformer-Based Deep Learning Models
- Legal Element-oriented Modeling with Multi-view Contrastive Learning for Legal Case Retrieval
- Study of positional encoding approaches for Audio Spectrogram Transformers
- CAPE: Encoding Relative Positions with Continuous Augmented Positional Embeddings
- MedGPT: Medical Concept Prediction from Clinical Narratives
- Distributed Deep Learning in Open Collaborations
- Is the Number of Trainable Parameters All That Actually Matters?
- Ripple Attention for Visual Perception with Sub-quadratic Complexity
- PermuteFormer: Efficient Relative Position Encoding for Long Sequences
- Conformer-based End-to-end Speech Recognition With Rotary Position Embedding
- Building a Foundation Model for Trajectory from Scratch
- Large Language Models -- the Future of Fundamental Physics?
- Out of Context: How important is Local Context in Neural Program Repair?