2 papers
cs.CL2026
news-crawler-LM: A Small Long-Context Model For High-Quality News Crawling
Pascal Stolzenburg, Jonas Golde, Max Dallabetta +1
Extracting structured content from news pages remains challenging due to heterogeneous HTML layouts, inconsistent markup, and substantial boilerplate such as navigation elements an…
cs.CL2024
Fundus: A Simple-to-Use News Scraper Optimized for High Quality Extractions
Max Dallabetta, Conrad Dobberstein, Adrian Breiding +1
This paper introduces Fundus, a user-friendly news scraper that enables users to obtain millions of high-quality news articles with just a few lines of code. Unlike existing news s…