information retrieval

Measuring What the Crawler Sees: Discovery Curves, Core Persistence, and Shell Dynamics in Longitudinal Web Crawls

arXiv:2607.13636

summary

The paper proposes a formal framework for analyzing longitudinal web crawls, introducing discovery curves and a two‑component urn model that separates a persistent core of URLs from a dynamic shell, and validates the approach on Common Crawl and the German Academic Web.

Abstract

A longitudinal web crawl is a sequence of partial samples of an evolving URL population. Pairwise containment between two crawls is the standard probe; under a simple \emph{urn} model of the crawl -- each round samples a fraction of the URLs and replaces a fraction -- it recovers two interpretable rates, per-round survival and coverage , but treats the population as uniform and consumes one pair at a time. In this work, we define a formal language for talking about a crawl. We extend this analysis with the \emph{discovery curve} , the cumulative URL footprint over a sliding window of crawls starting at , which under the same urn model is also a closed-form function of . Containment and the discovery curve are then two projections of one process: independent fits agree on when the urn is homogeneous, so any disagreement is itself a measurement. Applied to Common Crawl (2020--2025, domain granularity) and to the German Academic Web (GAW, URL granularity), the two projections disagree on both archives, and a two-component urn with a persistent core fraction alongside shell parameters reconciles the disagreement. A residual on remains, signaling that the shell itself is not homogeneous; is recorded as the scalar entry point to a rank-resolved generalization, which is left to follow-up work. \keywords{web archive \and crawl coverage \and discovery curve \and urn model \and two-component model \and URL lifetime}

16 pages, 4 figures, web metrics

Topics & keywords

#web crawling#longitudinal analysis#urn model#discovery curve#core persistence#shell dynamicsweb archivecrawl coveragediscovery curveurn modeltwo-component modelURL lifetime
Measuring What the Crawler Sees: Discovery Curves, Core Persistence, and Shell Dynamics in Longitudinal Web Crawls · wovepaper