Identifying Machine-Paraphrased Plagiarism
arXiv:2103.11909 · doi:10.1007/978-3-030-96957-8_34
Abstract
Employing paraphrasing tools to conceal plagiarized text is a severe threat to academic integrity. To enable the detection of machine-paraphrased text, we evaluate the effectiveness of five pre-trained word embedding models combined with machine-learning classifiers and eight state-of-the-art neural language models. We analyzed preprints of research papers, graduation theses, and Wikipedia articles, which we paraphrased using different configurations of the tools SpinBot and SpinnerChief. The best-performing technique, Longformer, achieved an average F1 score of 81.0% (F1=99.7% for SpinBot and F1=71.6% for SpinnerChief cases), while human evaluators achieved F1=78.4% for SpinBot and F1=65.6% for SpinnerChief cases. We show that the automated classification alleviates shortcomings of widely-used text-matching systems, such as Turnitin and PlagScan. To facilitate future research, all data, code, and two web applications showcasing our contributions are openly available at https://github.com/jpwahle/iconf22-paraphrase.
References in corpus (5)
- Distributed Representations of Sentences and Documents
- Neural Media Bias Detection Using Distant Supervision With BABE -- Bias Annotations By Experts
- Are Neural Language Models Good Plagiarists? A Benchmark for Neural Paraphrase Detection
- Improving Academic Plagiarism Detection for STEM Documents by Analyzing Mathematical Content and Citations
- Testing the Generalization of Neural Language Models for COVID-19 Misinformation Detection
Cited by in corpus (7)
- Are Neural Language Models Good Plagiarists? A Benchmark for Neural Paraphrase Detection
- How Large Language Models are Transforming Machine-Paraphrased Plagiarism
- Testing the Generalization of Neural Language Models for COVID-19 Misinformation Detection
- A Domain-adaptive Pre-training Approach for Language Bias Detection in News
- Paraphrase Types for Generation and Detection
- Analyzing Multi-Task Learning for Abstractive Text Summarization
- Discourse Features Enhance Detection of Document-Level Machine-Generated Content