6 papers
Agentic Clustering: Controllable Text Taxonomies via Multi-Agent Refinement
Simon Löwe, Emily Silcock
Recent text-clustering methods use large language models to propose a cluster taxonomy from a corpus and then assign each text to it. These pipelines are fundamentally programmatic…
News Deja Vu: Connecting Past and Present with Semantic Search
Brevin Franklin, Emily Silcock, Abhishek Arora +2
Social scientists and the general public often analyze contemporary events by drawing parallels with the past, a process complicated by the vast, noisy, and unstructured nature of…
Data Contamination Report from the 2024 CONDA Shared Task
Oscar Sainz, Iker GarcÃa-Ferrero, Alon Jacovi +25
The 1st Workshop on Data Contamination (CONDA 2024) focuses on all relevant aspects of data contamination in natural language processing, where data contamination is understood as…
Contrastive Entity Coreference and Disambiguation for Historical Texts
Abhishek Arora, Emily Silcock, Leander Heldring +1
Massive-scale historical document collections are crucial for social science research. Despite increasing digitization, these documents typically lack unique cross-document identif…
Newswire: A Large-Scale Structured Database of a Century of Historical News
Emily Silcock, Abhishek Arora, Luca D'Amico-Wong +1
In the U.S. historically, local newspapers drew their content largely from newswires like the Associated Press. Historians argue that newswires played a pivotal role in creating a…
Noise-Robust De-Duplication at Scale
Emily Silcock, Luca D'Amico-Wong, Jinglin Yang +1
Identifying near duplicates within large, noisy text corpora has a myriad of applications that range from de-duplicating training datasets, reducing privacy risk, and evaluating te…