works on

From the 2 of 6 linked papers with an AI index.

collaborators

6 papers

physics.soc-ph2026

Measuring What the Crawler Sees: Discovery Curves, Core Persistence, and Shell Dynamics in Longitudinal Web Crawls

Michael Paris, Hande Celikkanat, Luca Foppiano

The paper proposes a formal framework for analyzing longitudinal web crawls, introducing discovery curves and a two‑component urn model that separates a persistent core of URLs fro…

cs.DL2026

The INRIA DataLake: A Generic and Scalable Ecosystem of Pipelines for HAL Applied to Software Mentions Tracking

Luca Foppiano, Vipul Gupta, Samuel Scalbert +6

The paper describes the INRIA DataLake, a scalable ecosystem of interconnected pipelines that prepares scientific articles, extracts structured information such as software mention…

cs.CL2026

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94

Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…

cs.DL2026

Digging Up Citations: FOSSIL, a Dataset and Workflow for Reference Extraction in Law and the Humanities

Luca Foppiano, Christian Boulanger

Citation extraction tools are designed for the structured end-of-document bibliographies of the natural sciences, but law and humanities scholarship cites references primarily in f…

cs.CL2026

Construction of a Battery Research Knowledge Graph using a Global Open Catalog

Luca Foppiano, Sae Dieb, Malik Zain +3

Battery research is a rapidly growing and highly interdisciplinary field, making it increasingly difficult to track relevant expertise and identify potential collaborators across i…

cs.CL2026

SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing

Luca Foppiano, Sotaro Takeshita, Pedro Ortiz Suarez +6

SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English…