Attention in Natural Language Processing
arXiv:1902.02181 · doi:10.1109/TNNLS.2020.3019893
Abstract
Attention is an increasingly popular mechanism used in a wide range of neural architectures. The mechanism itself has been realized in a variety of formats. However, because of the fast-paced advances in this domain, a systematic overview of attention is still missing. In this article, we define a unified model for attention architectures in natural language processing, with a focus on those designed to work with vector representations of the textual data. We propose a taxonomy of attention models according to four dimensions: the representation of the input, the compatibility function, the distribution function, and the multiplicity of the input and/or output. We present the examples of how prior information can be exploited in attention models and discuss ongoing research efforts and open challenges in the area, providing the first extensive categorization of the vast body of literature in this exciting domain.
18 pages, 8 figures
References in corpus (16)
- A Structured Self-attentive Sentence Embedding
- Weight Uncertainty in Neural Networks
- Recurrent Models of Visual Attention
- Self-Normalizing Neural Networks
- Attention is not Explanation
- End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
- Grammar as a Foreign Language
- Structured Attention Networks
- Gated-Attention Architectures for Task-Oriented Language Grounding
- Attention Interpretability Across NLP Tasks
- Learning What Data to Learn
- Exploring Human-like Attention Supervision in Visual Question Answering
- Multi-focus Attention Network for Efficient Deep Reinforcement Learning
- Interactive Attention for Neural Machine Translation
- From Pixels to Objects: Cubic Visual Attention for Visual Question Answering
- A Question-Focused Multi-Factor Attention Network for Question Answering
Cited by in corpus (41)
- Attention Mechanism in Neural Networks: Where it Comes and Where it Goes
- Multimodal Data Integration for Oncology in the Era of Deep Neural Networks: A Review
- A Practical Survey on Faster and Lighter Transformers
- A Survey of Deep Learning Techniques for Neural Machine Translation
- Email Spam Detection Using Hierarchical Attention Hybrid Deep Learning Method
- An ensemble deep learning technique for detecting suicidal ideation from posts in social media platforms
- Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models
- Imbalance Knowledge-Driven Multi-modal Network for Land-Cover Semantic Segmentation Using Images and LiDAR Point Clouds
- A Survey of Visual Transformers
- Attention-gating for improved radio galaxy classification
- Attention U-Net as a surrogate model for groundwater prediction
- Evolving Modular Soft Robots without Explicit Inter-Module Communication using Local Self-Attention
- Neural Language Generation: Formulation, Methods, and Evaluation
- Forecasting GICs and geoelectric fields from solar wind data using LSTMs: application in Austria
- Exploring Transformers in Natural Language Generation: GPT, BERT, and XLNet
- An Argumentative Dialogue System for COVID-19 Vaccine Information
- Multi-Task Attentive Residual Networks for Argument Mining
- Do Machines and Humans Focus on Similar Code? Exploring Explainability of Large Language Models in Code Summarization
- A Survey of Text Representation Methods and Their Genealogy
- Machine Learning for Detecting Data Exfiltration: A Review
- AEI: Actors-Environment Interaction with Adaptive Attention for Temporal Action Proposals Generation
- Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?
- Thank you for Attention: A survey on Attention-based Artificial Neural Networks for Automatic Speech Recognition
- AMD-HookNet for Glacier Front Segmentation
- ProtoryNet - Interpretable Text Classification Via Prototype Trajectories
- Learning from similarity and information extraction from structured documents
- Self-Supervised Generative Models for Crystal Structures
- The Thousand Faces of Explainable AI Along the Machine Learning Life Cycle: Industrial Reality and Current State of Research
- Towards Interpreting Zoonotic Potential of Betacoronavirus Sequences With Attention
- Attribute-Guided Multi-Level Attention Network for Fine-Grained Fashion Retrieval
- MQTransformer: Multi-Horizon Forecasts with Context Dependent and Feedback-Aware Attention
- Clustering Text Using Attention
- An attention model to analyse the risk of agitation and urinary tract infections in people with dementia
- An Automatic Deep Learning Approach for Trailer Generation through Large Language Models
- Attention Models for Point Clouds in Deep Learning: A Survey
- Observing Dialogue in Therapy: Categorizing and Forecasting Behavioral Codes
- Tree-Constrained Graph Neural Networks For Argument Mining
- Spectraformer: A Unified Random Feature Framework for Transformer
- Interflow: Aggregating Multi-layer Feature Mappings with Attention Mechanism
- Interpretable by Design: Learning Predictors by Composing Interpretable Queries
- Text Classification with Lexicon from PreAttention Mechanism