6 citations · 8 across the 8 of their papers we have counts for
6 papers · 1 filter
Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics
Luca Foppiano
PDF corpora advertise their size in tokens but compute every rate they publish (coverage, OCR routing, re-fetch recovery, language mix) per document, and none decomposes its token…
Construction of a Battery Research Knowledge Graph using a Global Open Catalog
Luca Foppiano, Sae Dieb, Malik Zain +3
Battery research is a rapidly growing and highly interdisciplinary field, making it increasingly difficult to track relevant expertise and identify potential collaborators across i…
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…
SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing
Luca Foppiano, Sotaro Takeshita, Pedro Ortiz Suarez +6
SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English…
Mining experimental data from Materials Science literature with Large Language Models: an evaluation study
Luca Foppiano, Guillaume Lambard, Toshiyuki Amagasa +1
This study is dedicated to assessing the capabilities of large language models (LLMs) such as GPT-3.5-Turbo, GPT-4, and GPT-4-Turbo in extracting structured information from scient…
Semi-automatic staging area for high-quality structured data extraction from scientific literature
Luca Foppiano, Tomoya Mato, Kensei Terashima +7
We propose a semi-automatic staging area for efficiently building an accurate database of experimental physical properties of superconductors from literature, called SuperCon2, to…