activity
20232026
most citedSemi-automatic staging area for high-quality structured data extraction from scientific literature

6 citations · 8 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics

Luca Foppiano

PDF corpora advertise their size in tokens but compute every rate they publish (coverage, OCR routing, re-fetch recovery, language mix) per document, and none decomposes its token…

cs.CL2026

Construction of a Battery Research Knowledge Graph using a Global Open Catalog

Luca Foppiano, Sae Dieb, Malik Zain +3

Battery research is a rapidly growing and highly interdisciplinary field, making it increasingly difficult to track relevant expertise and identify potential collaborators across i…

cs.CL2026

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94

Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…

cs.CL20252 cited

SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing

Luca Foppiano, Sotaro Takeshita, Pedro Ortiz Suarez +6

SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English…

cs.CL2024

Mining experimental data from Materials Science literature with Large Language Models: an evaluation study

Luca Foppiano, Guillaume Lambard, Toshiyuki Amagasa +1

This study is dedicated to assessing the capabilities of large language models (LLMs) such as GPT-3.5-Turbo, GPT-4, and GPT-4-Turbo in extracting structured information from scient…

cs.CL20236 cited

Semi-automatic staging area for high-quality structured data extraction from scientific literature

Luca Foppiano, Tomoya Mato, Kensei Terashima +7

We propose a semi-automatic staging area for efficiently building an accurate database of experimental physical properties of superconductors from literature, called SuperCon2, to…