Longformer: The Long-Document Transformer
arXiv:2004.05150
Abstract
Transformer-based models are unable to process long sequences due to their self-attention operation, which scales quadratically with the sequence length. To address this limitation, we introduce the Longformer with an attention mechanism that scales linearly with sequence length, making it easy to process documents of thousands of tokens or longer. Longformer's attention mechanism is a drop-in replacement for the standard self-attention and combines a local windowed attention with a task motivated global attention. Following prior work on long-sequence transformers, we evaluate Longformer on character-level language modeling and achieve state-of-the-art results on text8 and enwik8. In contrast to most prior work, we also pretrain Longformer and finetune it on a variety of downstream tasks. Our pretrained Longformer consistently outperforms RoBERTa on long document tasks and sets new state-of-the-art results on WikiHop and TriviaQA. We finally introduce the Longformer-Encoder-Decoder (LED), a Longformer variant for supporting long document generative sequence-to-sequence tasks, and demonstrate its effectiveness on the arXiv summarization dataset.
Version 2 introduces the Longformer-Encoder-Decoder (LED) model
References in corpus (10)
- Sequence to Sequence Learning with Neural Networks
- WaveNet: A Generative Model for Raw Audio
- Generating Long Sequences with Sparse Transformers
- Reformer: The Efficient Transformer
- Big Bird: Transformers for Longer Sequences
- fairseq: A Fast, Extensible Toolkit for Sequence Modeling
- BP-Transformer: Modelling Long-Range Context via Binary Partitioning
- GMAT: Global Memory Augmentation for Transformers
- Select, Answer and Explain: Interpretable Multi-hop Reading Comprehension over Multiple Documents
- Graph Sequential Network for Reasoning over Sequences
Cited by in corpus (363)
- Survey of Hallucination in Natural Language Generation
- Transformers in Vision: A Survey
- On the Opportunities and Risks of Foundation Models
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- Linformer: Self-Attention with Linear Complexity
- Human Action Recognition from Various Data Modalities: A Review
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting
- DeBERTa: Decoding-enhanced BERT with Disentangled Attention
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- MetaFormer Baselines for Vision
- Dual Aspect Self-Attention based on Transformer for Remaining Useful Life Prediction
- Big Bird: Transformers for Longer Sequences
- Focal Self-attention for Local-Global Interactions in Vision Transformers
- TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale Up
- XCiT: Cross-Covariance Image Transformers
- Perceiver IO: A General Architecture for Structured Inputs & Outputs
- Spatio-Temporal Wind Speed Forecasting using Graph Networks and Novel Transformer Architectures
- Long Range Arena: A Benchmark for Efficient Transformers
- You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
- Video Transformers: A Survey
- The NLP Cookbook: Modern Recipes for Transformer based Deep Learning Architectures
- Differentiable Quantum Architecture Search
- A Comparative Study of Pretrained Language Models for Long Clinical Text
- Transformers and Large Language Models for Efficient Intrusion Detection Systems: A Comprehensive Survey
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- A Practical Survey on Faster and Lighter Transformers
- Making Pre-trained Language Models Better Few-shot Learners
- CenterCLIP: Token Clustering for Efficient Text-Video Retrieval
- Perceiver: General Perception with Iterative Attention
- Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer
- Random Feature Attention
- Rethinking Attention with Performers
- An Empirical Survey on Long Document Summarization: Datasets, Models and Metrics
- Fredformer: Frequency Debiased Transformer for Time Series Forecasting
- Pretrained Transformers for Text Ranking: BERT and Beyond
- Advances of Machine Learning in Materials Science: Ideas and Techniques
- Materials science in the era of large language models: a perspective
- Local-Global Context Aware Transformer for Language-Guided Video Segmentation
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
- Aligning AI With Shared Human Values
- Pretrained Transformers as Universal Computation Engines
- Rethinking Search: Making Domain Experts out of Dilettantes
- Word-Level Coreference Resolution
- Chasing Sparsity in Vision Transformers: An End-to-End Exploration
- Does the Magic of BERT Apply to Medical Code Assignment? A Quantitative Study
- Fastformer: Additive Attention Can Be All You Need
- Annotating Columns with Pre-trained Language Models
- Full Page Handwriting Recognition via Image to Sequence Extraction
- Towards Accurate Post-Training Quantization for Vision Transformer
- NeurIPS 2020 EfficientQA Competition: Systems, Analyses and Lessons Learned
- Neural Natural Language Processing for Long Texts: A Survey on Classification and Summarization
- Exploring the Synergies of Hybrid CNNs and ViTs Architectures for Computer Vision: A survey
- ConvTransformer: A Convolutional Transformer Network for Video Frame Synthesis
- Space-time Mixing Attention for Video Transformer
- VST++: Efficient and Stronger Visual Saliency Transformer
- Long-Short Transformer: Efficient Transformers for Language and Vision
- Discharge Summary Hospital Course Summarisation of In Patient Electronic Health Record Text with Clinical Concept Guided Deep Pre-Trained Transformer Models
- Hierarchical Label-wise Attention Transformer Model for Explainable ICD Coding
- Luna: Linear Unified Nested Attention
- Keeping Your Eye on the Ball: Trajectory Attention in Video Transformers
- LEXTREME: A Multi-Lingual and Multi-Task Benchmark for the Legal Domain
- LambdaNetworks: Modeling Long-Range Interactions Without Attention
- DeLighT: Deep and Light-weight Transformer
- Few-Shot Bot: Prompt-Based Learning for Dialogue Systems
- Self-Attention Between Datapoints: Going Beyond Individual Input-Output Pairs in Deep Learning
- GraphiT: Encoding Graph Structure in Transformers
- Hierarchical Learning for Generation with Long Source Sequences
- Identifying Machine-Paraphrased Plagiarism
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- A Dataset for Answering Time-Sensitive Questions
- A Unified Review of Deep Learning for Automated Medical Coding
- Language Models as Few-Shot Learner for Task-Oriented Dialogue Systems
- Mokey: Enabling Narrow Fixed-Point Inference for Out-of-the-Box Floating-Point Transformer Models
- Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
- SODA: A Natural Language Processing Package to Extract Social Determinants of Health for Cancer Studies
- Generation-Augmented Retrieval for Open-domain Question Answering
- HTLM: Hyper-Text Pre-Training and Prompting of Language Models
- MaX-DeepLab: End-to-End Panoptic Segmentation with Mask Transformers
- "This is Fake! Shared it by Mistake": Assessing the Intent of Fake News Spreaders
- Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
- Are Neural Language Models Good Plagiarists? A Benchmark for Neural Paraphrase Detection
- Text Guide: Improving the quality of long text classification by a text selection method based on feature importance
- Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention
- Domain-specific ChatBots for Science using Embeddings
- Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers
- How Large Language Models are Transforming Machine-Paraphrased Plagiarism
- ETC: Encoding Long and Structured Inputs in Transformers
- Adaptive Fourier Neural Operators: Efficient Token Mixers for Transformers
- Efficient Quantized Sparse Matrix Operations on Tensor Cores
- A Survey of Deep Learning Approaches for OCR and Document Understanding
- ConceptEVA: Concept-Based Interactive Exploration and Customization of Document Summaries
- GTrans: Grouping and Fusing Transformer Layers for Neural Machine Translation
- ZeroBERTo: Leveraging Zero-Shot Text Classification by Topic Modeling
- Mind the Gap: Assessing Temporal Generalization in Neural Language Models
- TVT: Transferable Vision Transformer for Unsupervised Domain Adaptation
- XAI in Computational Linguistics: Understanding Political Leanings in the Slovenian Parliament
- Language-Independent Sentiment Labelling with Distant Supervision: A Case Study for English, Sepedi and Setswana
- Vulgar Remarks Detection in Chittagonian Dialect of Bangla
- Detection of depression on social networks using transformers and ensembles
- Pile-Up Mitigation using Attention
- From Statistical Methods to Deep Learning, Automatic Keyphrase Prediction: A Survey
- GMAT: Global Memory Augmentation for Transformers
- Centroid Transformers: Learning to Abstract with Attention
- Match-Ignition: Plugging PageRank into Transformer for Long-form Text Matching
- RealFormer: Transformer Likes Residual Attention
- Summ^N: A Multi-Stage Summarization Framework for Long Input Dialogues and Documents
- Detecting Hallucinated Content in Conditional Neural Sequence Generation
- DeTiME: Diffusion-Enhanced Topic Modeling using Encoder-decoder based LLM
- Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval
- SYMBA: Symbolic Computation of Squared Amplitudes in High Energy Physics with Machine Learning
- Homogeneous Tokenizer Matters: Homogeneous Visual Tokenizer for Remote Sensing Image Understanding
- Transformers: "The End of History" for NLP?
- GNN-LM: Language Modeling based on Global Contexts via GNN
- Cluster-Former: Clustering-based Sparse Transformer for Long-Range Dependency Encoding
- Generating a Structured Summary of Numerous Academic Papers: Dataset and Method
- Coordination Among Neural Modules Through a Shared Global Workspace
- Efficient Transformers with Dynamic Token Pooling
- Conformer-Kernel with Query Term Independence for Document Retrieval
- Multitask Balanced and Recalibrated Network for Medical Code Prediction
- Multi-scale Time-stepping of Partial Differential Equations with Transformers
- Self-Supervised Document Similarity Ranking via Contextualized Language Models and Hierarchical Inference
- Bridging Textual and Tabular Data for Cross-Domain Text-to-SQL Semantic Parsing
- Pre-Trained Models: Past, Present and Future
- Lawformer: A Pre-trained Language Model for Chinese Legal Long Documents
- Towards Clinical Encounter Summarization: Learning to Compose Discharge Summaries from Prior Notes
- A Survey of Text Representation Methods and Their Genealogy
- The Power of Selecting Key Blocks with Local Pre-ranking for Long Document Information Retrieval
- Pre-Trained Language Models for Keyphrase Prediction: A Review
- Learning To Retrieve: How to Train a Dense Retrieval Model Effectively and Efficiently
- A Scoping Review of Natural Language Processing in Addressing Medically Inaccurate Information: Errors, Misinformation, and Hallucination
- Fast and Effective Biomedical Entity Linking Using a Dual Encoder
- Hi-BEHRT: Hierarchical Transformer-based model for accurate prediction of clinical events using multimodal longitudinal electronic health records
- Challenges in Domain-Specific Abstractive Summarization and How to Overcome them
- Fast Convergence of DETR with Spatially Modulated Co-Attention
- SparseBERT: Rethinking the Importance Analysis in Self-attention
- The NLP Task Effectiveness of Long-Range Transformers
- Graph-free Multi-hop Reading Comprehension: A Select-to-Guide Strategy
- Transformer Acceleration with Dynamic Sparse Attention
- Overview of the BioLaySumm 2023 Shared Task on Lay Summarization of Biomedical Research Articles
- DMDD: A Large-Scale Dataset for Dataset Mentions Detection
- SciCo: Hierarchical Cross-Document Coreference for Scientific Concepts
- Efficient Attentions for Long Document Summarization
- Conflicts, Villains, Resolutions: Towards models of Narrative Media Framing
- Abstractive Text Summarization Using the BRIO Training Paradigm
- A Meta-Evaluation of Faithfulness Metrics for Long-Form Hospital-Course Summarization
- ERNIE-Doc: A Retrospective Long-Document Modeling Transformer
- Learn To Remember: Transformer with Recurrent Memory for Document-Level Machine Translation
- Fourier Transformer: Fast Long Range Modeling by Removing Sequence Redundancy with FFT Operator
- Cura: Curation at Social Media Scale
- Connections are Expressive Enough: Universal Approximability of Sparse Transformers
- SC4ANM: Identifying Optimal Section Combinations for Automated Novelty Prediction in Academic Papers
- Intra-Document Cascading: Learning to Select Passages for Neural Document Ranking
- SFTformer: A Spatial-Frequency-Temporal Correlation-Decoupling Transformer for Radar Echo Extrapolation
- An Unbiased Transformer Source Code Learning with Semantic Vulnerability Graph
- Co-BERT: A Context-Aware BERT Retrieval Model Incorporating Local and Query-specific Context
- LazyFormer: Self Attention with Lazy Update
- Scatterbrain: Unifying Sparse and Low-rank Attention Approximation
- DynaPipe: Optimizing Multi-task Training through Dynamic Pipelines
- Gender Bias Detection in Court Decisions: A Brazilian Case Study
- Towards Suicide Prevention from Bipolar Disorder with Temporal Symptom-Aware Multitask Learning
- Personal Entity, Concept, and Named Entity Linking in Conversations
- Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search
- MS2: Multi-Document Summarization of Medical Studies
- Medical Knowledge Graph QA for Drug-Drug Interaction Prediction based on Multi-hop Machine Reading Comprehension
- Automated essay scoring using efficient transformer-based language models
- Towards a Multi-modal, Multi-task Learning based Pre-training Framework for Document Representation Learning
- Obstacle-Transformer: A Trajectory Prediction Network Based on Surrounding Trajectories
- SelfDoc: Self-Supervised Document Representation Learning
- LAMPRET: Layout-Aware Multimodal PreTraining for Document Understanding
- Kleister: Key Information Extraction Datasets Involving Long Documents with Complex Layouts
- Analysis of Twitter Users' Lifestyle Choices using Joint Embedding Model
- Multi-View Sequence-to-Sequence Models with Conversational Structure for Abstractive Dialogue Summarization
- End-to-End Attention-based Image Captioning
- Transformers for End-to-End InfoSec Tasks: A Feasibility Study
- Malicious or Benign? Towards Effective Content Moderation for Children's Videos
- SkillSpan: Hard and Soft Skill Extraction from English Job Postings
- TR-BERT: Dynamic Token Reduction for Accelerating BERT Inference
- Explaining The Efficacy of Counterfactually Augmented Data
- When FastText Pays Attention: Efficient Estimation of Word Representations using Constrained Positional Weighting
- Summarize, Outline, and Elaborate: Long-Text Generation via Hierarchical Supervision from Extractive Summaries
- Analysis of the Evolution of Advanced Transformer-Based Language Models: Experiments on Opinion Mining
- SJTU-NICT's Supervised and Unsupervised Neural Machine Translation Systems for the WMT20 News Translation Task
- Grid Search Hyperparameter Benchmarking of BERT, ALBERT, and LongFormer on DuoRC
- A Two-Phase Approach for Abstractive Podcast Summarization
- On the Trade-off between Redundancy and Local Coherence in Summarization
- Human Language Modeling
- Large Language Models as Autonomous Spacecraft Operators in Kerbal Space Program
- Learning the Simplicity of Scattering Amplitudes
- Connecting the Dots: What Graph-Based Text Representations Work Best for Text Classification Using Graph Neural Networks?
- D-FaST: Cognitive Signal Decoding with Disentangled Frequency-Spatial-Temporal Attention
- Deep Neural Networks and Brain Alignment: Brain Encoding and Decoding (Survey)
- Eigen Analysis of Self-Attention and its Reconstruction from Partial Computation
- Investigating the Effects of Sparse Attention on Cross-Encoders
- Longformer for MS MARCO Document Re-ranking Task
- FMMformer: Efficient and Flexible Transformer via Decomposed Near-field and Far-field Attention
- A Multilingual Translator to SQL with Database Schema Pruning to Improve Self-Attention
- Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion
- Not All Memories are Created Equal: Learning to Forget by Expiring
- Document Classification for COVID-19 Literature
- Question-Answering Approach to Evaluating Legal Summaries
- What Context Features Can Transformer Language Models Use?
- Code Summarization with Structure-induced Transformer
- Efficient Conformer with Prob-Sparse Attention Mechanism for End-to-EndSpeech Recognition
- Iterative Hierarchical Attention for Answering Complex Questions over Long Documents
- Language Modeling using LMUs: 10x Better Data Efficiency or Improved Scaling Compared to Transformers
- Which *BERT? A Survey Organizing Contextualized Encoders
- Multi-document Summarization: A Comparative Evaluation
- Journalistic Guidelines Aware News Image Captioning
- Complexity of Symbolic Representation in Working Memory of Transformer Correlates with the Complexity of a Task
- Around the GLOBE: Numerical Aggregation Question-Answering on Heterogeneous Genealogical Knowledge Graphs with Deep Neural Networks
- A Unified Efficient Pyramid Transformer for Semantic Segmentation
- Long-term series forecasting with Query Selector -- efficient model of sparse attention
- PoNet: Pooling Network for Efficient Token Mixing in Long Sequences
- KVT: k-NN Attention for Boosting Vision Transformers
- VidTr: Video Transformer Without Convolutions
- Natural Language Inference in Context -- Investigating Contextual Reasoning over Long Texts
- Can Transformers Reason About Effects of Actions?
- Multi-Scale 2D Temporal Adjacent Networks for Moment Localization with Natural Language
- Rissanen Data Analysis: Examining Dataset Characteristics via Description Length
- A Sliding-Window Approach to Automatic Creation of Meeting Minutes
- Adding Recurrence to Pretrained Transformers for Improved Efficiency and Context Size
- BUSTER: a "BUSiness Transaction Entity Recognition" dataset
- Conceptualizing Suicidal Behavior: Utilizing Explanations of Predicted Outcomes to Analyze Longitudinal Social Media Data
- DebateSum: A large-scale argument mining and summarization dataset
- Improved Transformer for High-Resolution GANs
- An Interpretable End-to-end Fine-tuning Approach for Long Clinical Text
- Hi-Transformer: Hierarchical Interactive Transformer for Efficient and Effective Long Document Modeling
- Staircase Attention for Recurrent Processing of Sequences
- Classifying Long Clinical Documents with Pre-trained Transformers
- ReadTwice: Reading Very Large Documents with Memories
- Transfer Learning for Multi-lingual Tasks -- a Survey
- baller2vec++: A Look-Ahead Multi-Entity Transformer For Modeling Coordinated Agents
- Optimizing small BERTs trained for German NER
- Layered gradient accumulation and modular pipeline parallelism: fast and efficient training of large language models
- Open Question Answering over Tables and Text
- Video Transformer for Deepfake Detection with Incremental Learning
- TANGNN: a Concise, Scalable and Effective Graph Neural Networks with Top-m Attention Mechanism for Graph Representation Learning
- LongKey: Keyphrase Extraction for Long Documents
- GLiT: Neural Architecture Search for Global and Local Image Transformer
- Long-Range Transformer Architectures for Document Understanding
- G-Transformer for Document-level Machine Translation
- Unsupervised Document Embedding via Contrastive Augmentation
- Multi-Task Prediction of Clinical Outcomes in the Intensive Care Unit using Flexible Multimodal Transformers
- LUKE-Graph: A Transformer-based Approach with Gated Relational Graph Attention for Cloze-style Reading Comprehension
- Is In-hospital Meta-information Useful for Abstractive Discharge Summary Generation?
- Transformation Invariant Cancerous Tissue Classification Using Spatially Transformed DenseNet
- Deep Insights into Cognitive Decline: A Survey of Leveraging Non-Intrusive Modalities with Deep Learning Techniques
- PGT: Pseudo Relevance Feedback Using a Graph-Based Transformer
- Enhanced Labeling Technique for Reddit Text and Fine-Tuned Longformer Models for Classifying Depression Severity in English and Luganda
- On Generating Extended Summaries of Long Documents
- Investigating the detection of Tortured Phrases in Scientific Literature
- MMCoVaR: Multimodal COVID-19 Vaccine Focused Data Repository for Fake News Detection and a Baseline Architecture for Classification
- N-ODE Transformer: A Depth-Adaptive Variant of the Transformer Using Neural Ordinary Differential Equations
- Multi-Scale Local-Temporal Similarity Fusion for Continuous Sign Language Recognition
- Long Document Ranking with Query-Directed Sparse Transformer
- EasyTransfer -- A Simple and Scalable Deep Transfer Learning Platform for NLP Applications
- Current Limitations of Language Models: What You Need is Retrieval
- Transformer-Based Behavioral Representation Learning Enables Transfer Learning for Mobile Sensing in Small Datasets
- Training Transformers for Information Security Tasks: A Case Study on Malicious URL Prediction
- DoSSIER@COLIEE 2021: Leveraging dense retrieval and summarization-based re-ranking for case law retrieval
- Self-supervised Answer Retrieval on Clinical Notes
- Sparse Factorization of Large Square Matrices
- On Learning the Transformer Kernel
- On the Distribution, Sparsity, and Inference-time Quantization of Attention Values in Transformers
- The Role of Global and Local Context in Named Entity Recognition
- Automating Chapter-Level Classification for Electronic Theses and Dissertations
- Predicting COVID-19 Patient Shielding: A Comprehensive Study
- Global memory transformer for processing long documents
- DocNLI: A Large-scale Dataset for Document-level Natural Language Inference
- Attention Mechanism with Energy-Friendly Operations
- TransMask: A Compact and Fast Speech Separation Model Based on Transformer
- Adaptive Semiparametric Language Models
- Shatter: An Efficient Transformer Encoder with Single-Headed Self-Attention and Relative Sequence Partitioning
- Dynamic Sliding Window for Meeting Summarization
- Smart Bird: Learnable Sparse Attention for Efficient and Effective Transformer
- EL-Attention: Memory Efficient Lossless Attention for Generation
- SMYRF: Efficient Attention using Asymmetric Clustering
- Pre-training Protein Language Models with Label-Agnostic Binding Pairs Enhances Performance in Downstream Tasks
- Weakly-supervised Text Classification Based on Keyword Graph
- The DEformer: An Order-Agnostic Distribution Estimating Transformer
- FastSeq: Make Sequence Generation Faster
- TrackFormers: In Search of Transformer-Based Particle Tracking for the High-Luminosity LHC Era
- Interpretable Self-supervised Multi-task Learning for COVID-19 Information Retrieval and Extraction
- H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for Sequences
- MORTY: Structured Summarization for Targeted Information Extraction from Scholarly Articles
- Token Pooling in Vision Transformers
- StreaMulT: Streaming Multimodal Transformer for Heterogeneous and Arbitrary Long Sequential Data
- -former: Infinite Memory Transformer
- Efficient Domain Adaptation of Language Models via Adaptive Tokenization
- GroupLink: An End-to-end Multitask Method for Word Grouping and Relation Extraction in Form Understanding
- Can Deep Neural Networks Predict Data Correlations from Column Names?
- Do Long-Range Language Models Actually Use Long-Range Context?
- Multi-Domain Transformer-Based Counterfactual Augmentation for Earnings Call Analysis
- Medical Code Prediction from Discharge Summary: Document to Sequence BERT using Sequence Attention
- Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
- Natural Language Processing to Detect Cognitive Concerns in Electronic Health Records Using Deep Learning
- When Attention Meets Fast Recurrence: Training Language Models with Reduced Compute
- Linear Self-Attention Approximation via Trainable Feedforward Kernel
- Beyond Nyströmformer -- Approximation of self-attention by Spectral Shifting
- Simple Hack for Transformers against Heavy Long-Text Classification on a Time- and Memory-Limited GPU Service
- Pre-trained Language Model based Ranking in Baidu Search
- Relation/Entity-Centric Reading Comprehension
- PLSUM: Generating PT-BR Wikipedia by Summarizing Multiple Websites
- Skyformer: Remodel Self-Attention with Gaussian Kernel and Nyström Method
- Domain Agnostic Few-Shot Learning For Document Intelligence
- ProSTformer: Pre-trained Progressive Space-Time Self-attention Model for Traffic Flow Forecasting
- Exploiting a Zoo of Checkpoints for Unseen Tasks
- DSBERT:Unsupervised Dialogue Structure learning with BERT
- A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers
- T-EMDE: Sketching-based global similarity for cross-modal retrieval
- Searching Personal Collections
- Decoupled Transformer for Scalable Inference in Open-domain Question Answering
- Understanding and Overcoming the Challenges of Efficient Transformer Quantization
- Description-based Label Attention Classifier for Explainable ICD-9 Classification
- What Makes a Star Teacher? A Hierarchical BERT Model for Evaluating Teacher's Performance in Online Education
- Improving Patent Mining and Relevance Classification using Transformers
- Don't be Contradicted with Anything! CI-ToD: Towards Benchmarking Consistency for Task-oriented Dialogue System
- MiRANews: Dataset and Benchmarks for Multi-Resource-Assisted News Summarization
- Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax Hierarchy
- EfficientBERT: Progressively Searching Multilayer Perceptron via Warm-up Knowledge Distillation
- Transformer Models for Text Coherence Assessment
- Data-Informed Global Sparseness in Attention Mechanisms for Deep Neural Networks
- Does Dialog Length matter for Next Response Selection task? An Empirical Study
- No Answer is Better Than Wrong Answer: A Reflection Model for Document Level Machine Reading Comprehension
- RoR: Read-over-Read for Long Document Machine Reading Comprehension
- StreamHover: Livestream Transcript Summarization and Annotation
- Speech Summarization using Restricted Self-Attention
- VAULT: VAriable Unified Long Text Representation for Machine Reading Comprehension
- SummerTime: Text Summarization Toolkit for Non-experts
- An Exploratory Study on Long Dialogue Summarization: What Works and What's Next
- Can Transformer Models Measure Coherence In Text? Re-Thinking the Shuffle Test
- Graph Convolutional Network for Swahili News Classification
- DIALKI: Knowledge Identification in Conversational Systems through Dialogue-Document Contextualization
- Knowledge Enhanced Sports Game Summarization
- Learning Opinion Summarizers by Selecting Informative Reviews
- Effective Distributed Representations for Academic Expert Search
- Corruption Is Not All Bad: Incorporating Discourse Structure into Pre-training via Corruption for Essay Scoring
- Efficient Transformer for Direct Speech Translation
- PermuteFormer: Efficient Relative Position Encoding for Long Sequences
- Vector-Vector-Matrix Architecture: A Novel Hardware-Aware Framework for Low-Latency Inference in NLP Applications
- Highly Parallel Autoregressive Entity Linking with Discriminative Correction
- Revisiting Simple Neural Probabilistic Language Models
- Simulated Chats for Building Dialog Systems: Learning to Generate Conversations from Instructions
- Semantic Frame Forecast
- ReadOnce Transformers: Reusable Representations of Text for Transformers
- Viola: A Topic Agnostic Generate-and-Rank Dialogue System
- Determining the Credibility of Science Communication
- Plot-guided Adversarial Example Construction for Evaluating Open-domain Story Generation
- Corpus-level and Concept-based Explanations for Interpretable Document Classification
- Go Forth and Prosper: Language Modeling with Ancient Textual History
- On Inductive Biases for Machine Learning in Data Constrained Settings
- Learning Hard Retrieval Decoder Attention for Transformers
- Transfer training from smaller language model
- A Hierarchical Neural Framework for Classification and its Explanation in Large Unstructured Legal Documents
- A Neural Edge-Editing Approach for Document-Level Relation Graph Extraction
- Input-independent Attention Weights Are Expressive Enough: A Study of Attention in Self-supervised Audio Transformers
- Hierarchical Context-Aware Transformers for Non-Autoregressive Text to Speech
- HETFORMER: Heterogeneous Transformer with Sparse Attention for Long-Text Extractive Summarization
- FarFetched: Entity-centric Reasoning and Claim Validation for the Greek Language based on Textually Represented Environments
- Contrastive Document Representation Learning with Graph Attention Networks
- DCT: Dynamic Compressive Transformer for Modeling Unbounded Sequence
- SciXGen: A Scientific Paper Dataset for Context-Aware Text Generation
- BumbleBee: A Transformer for Music