Beyond 512 Tokens: Siamese Multi-depth Transformer-based Hierarchical Encoder for Long-Form Document Matching
arXiv:2004.12297 · doi:10.1145/3340531.3411908
Abstract
Many natural language processing and information retrieval problems can be formalized as the task of semantic matching. Existing work in this area has been largely focused on matching between short texts (e.g., question answering), or between a short and a long text (e.g., ad-hoc retrieval). Semantic matching between long-form documents, which has many important applications like news recommendation, related article recommendation and document clustering, is relatively less explored and needs more research effort. In recent years, self-attention based models like Transformers and BERT have achieved state-of-the-art performance in the task of text matching. These models, however, are still limited to short text like a few sentences or one paragraph due to the quadratic computational complexity of self-attention with respect to input text length. In this paper, we address the issue by proposing the Siamese Multi-depth Transformer-based Hierarchical (SMITH) Encoder for long-form document matching. Our model contains several innovations to adapt self-attention models for longer text input. In order to better capture sentence level semantic relations within a document, we pre-train the model with a novel masked sentence block language modeling task in addition to the masked word language modeling task used by BERT. Our experimental results on several benchmark datasets for long-form document matching show that our proposed SMITH model outperforms the previous state-of-the-art models including hierarchical attention, multi-depth attention-based hierarchical recurrent neural network, and BERT. Comparing to BERT based baselines, our model is able to increase maximum input text length from 512 to 2048. We will open source a Wikipedia based benchmark dataset, code and a pre-trained checkpoint to accelerate future research on long-form document matching.
Accepted as a full paper in CIKM 2020
References in corpus (13)
- Convolutional Neural Network Architectures for Matching Natural Language Sentences
- A Deep Relevance Matching Model for Ad-hoc Retrieval
- End-to-End Neural Ad-hoc Ranking with Kernel Pooling
- Generating Long Sequences with Sparse Transformers
- Deeper Text Understanding for IR with Contextual Neural Language Modeling
- Axial Attention in Multidimensional Transformers
- Reformer: The Efficient Transformer
- Overview of the TREC 2020 deep learning track
- Overview of the TREC 2019 deep learning track
- HIBERT: Document Level Pre-training of Hierarchical Bidirectional Transformers for Document Summarization
- Compressive Transformers for Long-Range Sequence Modelling
- Blockwise Self-Attention for Long Document Understanding
- A Deep Look into Neural Ranking Models for Information Retrieval
Cited by in corpus (14)
- Long Range Arena: A Benchmark for Efficient Transformers
- Hierarchical Transformer with Spatio-Temporal Context Aggregation for Next Point-of-Interest Recommendation
- Match-Ignition: Plugging PageRank into Transformer for Long-form Text Matching
- Machine Learning for Violence Risk Assessment Using Dutch Clinical Notes
- Self-Supervised Document Similarity Ranking via Contextualized Language Models and Hierarchical Inference
- Legal Element-oriented Modeling with Multi-view Contrastive Learning for Legal Case Retrieval
- Neural Models for Offensive Language Detection
- Hi-Transformer: Hierarchical Interactive Transformer for Efficient and Effective Long Document Modeling
- Can Deep Neural Networks Predict Data Correlations from Column Names?
- Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
- RoR: Read-over-Read for Long Document Machine Reading Comprehension
- Contrastive Document Representation Learning with Graph Attention Networks
- Multilevel Text Alignment with Cross-Document Attention
- Paragraph-level Rationale Extraction through Regularization: A case study on European Court of Human Rights Cases