most citedAmerican Stories: A Large-Scale Structured Text Dataset of Historical U.S. Newspapers

7 citations · 7 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2024

Data Contamination Report from the 2024 CONDA Shared Task

Oscar Sainz, Iker García-Ferrero, Alon Jacovi +25

The 1st Workshop on Data Contamination (CONDA 2024) focuses on all relevant aspects of data contamination in natural language processing, where data contamination is understood as…

cs.CL2024

Newswire: A Large-Scale Structured Database of a Century of Historical News

Emily Silcock, Abhishek Arora, Luca D'Amico-Wong +1

In the U.S. historically, local newspapers drew their content largely from newswires like the Associated Press. Historians argue that newswires played a pivotal role in creating a…

cs.GT2024

Disrupting Bipartite Trading Networks: Matching for Revenue Maximization

Luca D'Amico-Wong, Yannai A. Gonczarowski, Gary Qiurui Ma +1

We model the role of an online platform disrupting a market with unit-demand buyers and unit-supply sellers. Each seller can transact with a subset of the buyers whom she already k…

cs.LG2024

Easy as ABCs: Unifying Boltzmann Q-Learning and Counterfactual Regret Minimization

Luca D'Amico-Wong, Hugh Zhang, Marc Lanctot +1

We propose ABCs (Adaptive Branching through Child stationarity), a best-of-both-worlds algorithm combining Boltzmann Q-learning (BQL), a classic reinforcement learning algorithm fo…

cs.CL20237 cited

American Stories: A Large-Scale Structured Text Dataset of Historical U.S. Newspapers

Melissa Dell, Jacob Carlson, Tom Bryan +7

Existing full text datasets of U.S. public domain newspapers do not recognize the often complex layouts of newspaper scans, and as a result the digitized content scrambles texts fr…