Comparing the writing style of real and artificial papers
arXiv:1506.05702 · doi:10.1007/s11192-015-1637-z
Abstract
Recent years have witnessed the increase of competition in science. While promoting the quality of research in many cases, an intense competition among scientists can also trigger unethical scientific behaviors. To increase the total number of published papers, some authors even resort to software tools that are able to produce grammatical, but meaningless scientific manuscripts. Because automatically generated papers can be misunderstood as real papers, it becomes of paramount importance to develop means to identify these scientific frauds. In this paper, I devise a methodology to distinguish real manuscripts from those generated with SCIGen, an automatic paper generator. Upon modeling texts as complex networks (CN), it was possible to discriminate real from fake papers with at least 89\% of accuracy. A systematic analysis of features relevance revealed that the accessibility and betweenness were useful in particular cases, even though the relevance depended upon the dataset. The successful application of the methods described here show, as a proof of principle, that network features can be used to identify scientific gibberish papers. In addition, the CN-based approach can be combined in a straightforward fashion with traditional statistical language processing methods to improve the performance in identifying artificially generated papers.
To appear in Scientometrics (2015)
References in corpus (9)
- Finding community structure in networks using the eigenvectors of matrices
- Diffusion of scientific credits and the ranking of scientists
- Probing the topological properties of complex networks modeling short written texts
- Patterns of Text Reuse in a Scientific Corpus
- Wikipedia information flow analysis reveals the scale-free architecture of the Semantic Space
- Word sense disambiguation via high order of learning in complex networks
- Identification of Literary Movements Using Complex Networks to Represent Texts
- Complex networks analysis of language complexity
- Algorithmic Detection of Computer Generated Text
Cited by in corpus (22)
- Random Walks: A Review of Algorithms and Applications
- Decoding ChatGPT: A Taxonomy of Existing Research, Current Challenges, and Possible Future Directions
- Using network science and text analytics to produce surveys in a scientific topic
- Word sense disambiguation: a complex network approach
- Extractive Multi-document Summarization Using Multilayer Networks
- Paragraph-based complex networks: application to document classification and authenticity verification
- Exploratory analysis of text duplication in peer-review reveals peer-review fraud and paper mills
- Word sense induction using word embeddings and community detection in complex networks
- Automated scholarly paper review: Concepts, technologies, and challenges
- A complex network approach to political analysis: application to the Brazilian Chamber of Deputies
- Topic segmentation via community detection in complex networks
- Connecting Network Science and Information Theory
- The Dynamics of Knowledge Acquisition via Self-Learning in Complex Networks
- Semantic flow in language networks
- Analyzing the relationship between text features and research proposal productivity
- Complex systems: features, similarity and connectivity
- Quantitative Discourse Cohesion Analysis of Scientific Scholarly Texts using Multilayer Networks
- Classical and Quantum Random Walks to Identify Leaders in Criminal Networks
- Language Networks: a Practical Approach
- K-method of calculating the mutual influence of nodes in a directed weight complex networks
- Classification of abrupt changes along viewing profiles of scientific articles
- Accessibility: A Generalization of the Node Degree (A Tutorial)