3 papers
cs.CL2020
Quasi Error-free Text Classification and Authorship Recognition in a large Corpus of English Literature based on a Novel Feature Set
Arthur M. Jacobs, Annette Kinder
The Gutenberg Literary English Corpus (GLEC) provides a rich source of textual data for research in digital humanities, computational linguistics or neurocognitive poetics. However…
cs.CL2018
Features of word similarity
Arthur M. Jacobs, Annette Kinder
In this theoretical note we compare different types of computational models of word similarity and association in their ability to predict a set of about 900 rating data. Using reg…
cs.CL2018
Explorations in an English Poetry Corpus: A Neurocognitive Poetics Perspective
Arthur M. Jacobs
This paper describes a corpus of about 3000 English literary texts with about 250 million words extracted from the Gutenberg project that span a range of genres from both fiction a…