54 citations · 153 across the 6 of their papers we have counts for
8 papers
Multilayer Networks for Text Analysis with Multiple Data Types
Charles C. Hyland, Yuanming Tao, Lamiae Azizi +3
We are interested in the widespread problem of clustering documents and finding topics in large collections of written documents in the presence of metadata and hyperlinks. To tack…
A Multilingual Entity Linking System for Wikipedia with a Machine-in-the-Loop Approach
Martin Gerlach, Marshall Miller, Rita Ho +2
Hyperlinks constitute the backbone of the Web; they enable user navigation, information discovery, content ranking, and many other crucial services on the Internet. In particular,…
Language-agnostic Topic Classification for Wikipedia
Isaac Johnson, Martin Gerlach, Diego Sáez-Trumper
A major challenge for many analyses of Wikipedia dynamics -- e.g., imbalances in content quality, geographic differences in what content is popular, what types of articles attract…
A Taxonomy of Knowledge Gaps for Wikimedia Projects (Second Draft)
Miriam Redi, Martin Gerlach, Isaac Johnson +2
In January 2019, prompted by the Wikimedia Movement's 2030 strategic direction, the Research team at the Wikimedia Foundation identified the need to develop a knowledge gaps index…
Testing statistical laws in complex systems
Martin Gerlach, Eduardo G. Altmann
The availability of large datasets requires an improved view on statistical laws in complex systems, such as Zipf's law of word frequencies, the Gutenberg-Richter law of earthquake…
A new evaluation framework for topic modeling algorithms based on synthetic corpora
Hanyu Shi, Martin Gerlach, Isabel Diersen +2
Topic models are in widespread use in natural language processing and beyond. Here, we propose a new framework for the evaluation of probabilistic topic modeling algorithms based o…