PatternRank: Leveraging Pretrained Language Models and Part of Speech for Unsupervised Keyphrase Extraction
arXiv:2210.05245 · doi:10.5220/0011546600003335
Abstract
Keyphrase extraction is the process of automatically selecting a small set of most relevant phrases from a given text. Supervised keyphrase extraction approaches need large amounts of labeled training data and perform poorly outside the domain of the training data. In this paper, we present PatternRank, which leverages pretrained language models and part-of-speech for unsupervised keyphrase extraction from single documents. Our experiments show PatternRank achieves higher precision, recall and F1-scores than previous state-of-the-art approaches. In addition, we present the KeyphraseVectorizers package, which allows easy modification of part-of-speech patterns for candidate keyphrase selection, and hence adaptation of our approach to any domain.
Accepted to 14th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management - KDIR
References in corpus (2)
Cited by in corpus (7)
- Evaluating Unsupervised Text Classification: Zero-shot and Similarity-based Approaches
- Pre-Trained Language Models for Keyphrase Prediction: A Review
- AspectCSE: Sentence Embeddings for Aspect-based Semantic Textual Similarity Using Contrastive Learning and Structured Knowledge
- Efficient Domain Adaptation of Sentence Embeddings Using Adapters
- LongKey: Keyphrase Extraction for Long Documents
- Unsupervised extraction of local and global keywords from a single text
- Dynamik: Syntactically-Driven Dynamic Font Sizing for Emphasis of Key Information