From the 2 of 6 linked papers with an AI index.
6 papers
Measuring What the Crawler Sees: Discovery Curves, Core Persistence, and Shell Dynamics in Longitudinal Web Crawls
Michael Paris, Hande Celikkanat, Luca Foppiano
The paper proposes a formal framework for analyzing longitudinal web crawls, introducing discovery curves and a two‑component urn model that separates a persistent core of URLs fro…
The INRIA DataLake: A Generic and Scalable Ecosystem of Pipelines for HAL Applied to Software Mentions Tracking
Luca Foppiano, Vipul Gupta, Samuel Scalbert +6
The paper describes the INRIA DataLake, a scalable ecosystem of interconnected pipelines that prepares scientific articles, extracts structured information such as software mention…
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…
Digging Up Citations: FOSSIL, a Dataset and Workflow for Reference Extraction in Law and the Humanities
Luca Foppiano, Christian Boulanger
Citation extraction tools are designed for the structured end-of-document bibliographies of the natural sciences, but law and humanities scholarship cites references primarily in f…
Construction of a Battery Research Knowledge Graph using a Global Open Catalog
Luca Foppiano, Sae Dieb, Malik Zain +3
Battery research is a rapidly growing and highly interdisciplinary field, making it increasingly difficult to track relevant expertise and identify potential collaborators across i…
SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing
Luca Foppiano, Sotaro Takeshita, Pedro Ortiz Suarez +6
SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English…