activity
20102023
most citedPOLYGLOT-NER: Massive Multilingual Named Entity Recognition

38 citations · 63 across the 11 of their papers we have counts for

collaborators

11 papers

cs.CL2023

Analyzing Film Adaptation through Narrative Alignment

Tanzir Pial, Shahreen Salim, Charuta Pethe +2

Novels are often adapted into feature films, but the differences between the two media usually require dropping sections of the source text from the movie script. Here we study thi…

cs.CL2023

GNAT: A General Narrative Alignment Tool

Tanzir Pial, Steven Skiena

Algorithmic sequence alignment identifies similar segments shared between pairs of documents, and is fundamental to many NLP tasks. But it is difficult to recognize similarities be…

cs.CL20231 cited

STONYBOOK: A System and Resource for Large-Scale Analysis of Novels

Charuta Pethe, Allen Kim, Rajesh Prabhakar +2

Books have historically been the primary mechanism through which narratives are transmitted. We have developed a collection of resources for the large-scale analysis of novels, inc…

cs.DS2023

Accelerating Personalized PageRank Vector Computation

Zhen Chen, Xingzhi Guo, Baojian Zhou +2

Personalized PageRank Vectors are widely used as fundamental graph-learning tools for detecting anomalous spammers, learning graph embeddings, and training graph neural networks. T…

cs.CG2023

Does it pay to optimize AUC?

Baojian Zhou, Steven Skiena

The Area Under the ROC Curve (AUC) is an important model metric for evaluating binary classifiers, and many algorithms have been proposed to optimize AUC approximately. It raises t…

cs.CL201615 cited

False-Friend Detection and Entity Matching via Unsupervised Transliteration

Yanqing Chen, Steven Skiena

Transliterations play an important role in multilingual entity reference resolution, because proper names increasingly travel between languages in news and social media. Previous w…