Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Quantifying Geospatial in the Common Crawl Corpus
Ilya Ilyankou, Meihui Wang, Stefano Cavazzi +1
Large language models (LLMs) exhibit emerging geospatial capabilities, stemming from their pre-training on vast unlabelled text datasets that are often derived from the Common Craw…
cs.CL2024
CC-GPX: Extracting High-Quality Annotated Geospatial Data from Common Crawl
Ilya Ilyankou, Meihui Wang, Stefano Cavazzi +1
The Common Crawl (CC) corpus is the largest open web crawl dataset containing 9.5+ petabytes of data captured since 2008. The dataset is instrumental in training large language mod…