2 papers
cs.IR2018
iLCM - A Virtual Research Infrastructure for Large-Scale Qualitative Data
Andreas Niekler, Arnim Bleier, Christian Kahmann +5
The iLCM project pursues the development of an integrated research environment for the analysis of structured and unstructured data in a "Software as a Service" architecture (SaaS)…
cs.CL2017
TokTrack: A Complete Token Provenance and Change Tracking Dataset for the English Wikipedia
Fabian Flöck, Kenan Erdogan, Maribel Acosta
We present a dataset that contains every instance of all tokens (~ words) ever written in undeleted, non-redirect English Wikipedia articles until October 2016, in total 13,545,349…