collaborators

6 papers

cs.CL2026

Agentic Clustering: Controllable Text Taxonomies via Multi-Agent Refinement

Simon Löwe, Emily Silcock

Recent text-clustering methods use large language models to propose a cluster taxonomy from a corpus and then assign each text to it. These pipelines are fundamentally programmatic…

cs.CL2024

News Deja Vu: Connecting Past and Present with Semantic Search

Brevin Franklin, Emily Silcock, Abhishek Arora +2

Social scientists and the general public often analyze contemporary events by drawing parallels with the past, a process complicated by the vast, noisy, and unstructured nature of…

cs.CL2024

Data Contamination Report from the 2024 CONDA Shared Task

Oscar Sainz, Iker García-Ferrero, Alon Jacovi +25

The 1st Workshop on Data Contamination (CONDA 2024) focuses on all relevant aspects of data contamination in natural language processing, where data contamination is understood as…

cs.CL2024

Contrastive Entity Coreference and Disambiguation for Historical Texts

Abhishek Arora, Emily Silcock, Leander Heldring +1

Massive-scale historical document collections are crucial for social science research. Despite increasing digitization, these documents typically lack unique cross-document identif…

cs.CL2024

Newswire: A Large-Scale Structured Database of a Century of Historical News

Emily Silcock, Abhishek Arora, Luca D'Amico-Wong +1

In the U.S. historically, local newspapers drew their content largely from newswires like the Associated Press. Historians argue that newswires played a pivotal role in creating a…

cs.CL2024

Noise-Robust De-Duplication at Scale

Emily Silcock, Luca D'Amico-Wong, Jinglin Yang +1

Identifying near duplicates within large, noisy text corpora has a myriad of applications that range from de-duplicating training datasets, reducing privacy risk, and evaluating te…