Pre-Trained Language Models for Keyphrase Prediction: A Review
arXiv:2409.01087 · doi:10.1016/j.icte.2024.05.015
Abstract
Keyphrase Prediction (KP) is essential for identifying keyphrases in a document that can summarize its content. However, recent Natural Language Processing (NLP) advances have developed more efficient KP models using deep learning techniques. The limitation of a comprehensive exploration jointly both keyphrase extraction and generation using pre-trained language models spotlights a critical gap in the literature, compelling our survey paper to bridge this deficiency and offer a unified and in-depth analysis to address limitations in previous surveys. This paper extensively examines the topic of pre-trained language models for keyphrase prediction (PLM-KP), which are trained on large text corpora via different learning (supervisor, unsupervised, semi-supervised, and self-supervised) techniques, to provide respective insights into these two types of tasks in NLP, precisely, Keyphrase Extraction (KPE) and Keyphrase Generation (KPG). We introduce appropriate taxonomies for PLM-KPE and KPG to highlight these two main tasks of NLP. Moreover, we point out some promising future directions for predicting keyphrases.
References in corpus (18)
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Linformer: Self-Attention with Linear Complexity
- Publicly Available Clinical BERT Embeddings
- Learning to Extract Keyphrases from Text
- Complex Network based Supervised Keyword Extractor
- Keyphrase Generation for Scientific Document Retrieval
- Arabic Keyphrase Extraction using Linguistic knowledge and Machine Learning Techniques
- DivGraphPointer: A Graph Pointer Network for Extracting Diverse Keyphrases
- Keyphrase Extraction from Scholarly Articles as Sequence Labeling using Contextualized Embeddings
- From Statistical Methods to Deep Learning, Automatic Keyphrase Prediction: A Survey
- Ring Attention with Blockwise Transformers for Near-Infinite Context
- Keyphrase Prediction With Pre-trained Language Model
- Applying a Generic Sequence-to-Sequence Model for Simple and Effective Keyphrase Generation
- Unsupervised Key-phrase Extraction and Clustering for Classification Scheme in Scientific Publications
- GLEAKE: Global and Local Embedding Automatic Keyphrase Extraction
- LDKP: A Dataset for Identifying Keyphrases from Long Scientific Documents
- Cross-Media Keyphrase Prediction: A Unified Framework with Multi-Modality Multi-Head Attention and Image Wordings
- On Leveraging Encoder-only Pre-trained Language Models for Effective Keyphrase Generation