Efficient Attentions for Long Document Summarization
arXiv:2104.02112
Abstract
The quadratic computational and memory complexities of large Transformers have limited their scalability for long document summarization. In this paper, we propose Hepos, a novel efficient encoder-decoder attention with head-wise positional strides to effectively pinpoint salient information from the source. We further conduct a systematic study of existing efficient self-attentions. Combined with Hepos, we are able to process ten times more tokens than existing models that use full attentions. For evaluation, we present a new dataset, GovReport, with significantly longer documents and summaries. Results show that our models produce significantly higher ROUGE scores than competitive comparisons, including new state-of-the-art results on PubMed. Human evaluation also shows that our models generate more informative summaries with fewer unfaithful errors.
Accepted at NAACL 2021 as a long paper
References in corpus (7)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Generating Long Sequences with Sparse Transformers
- Reformer: The Efficient Transformer
- Sparse Sinkhorn Attention
- Scientific Paper Summarization Using Citation Summary Networks
- SEAL: Segment-wise Extractive-Abstractive Long-form Text Summarization