Automatic extraction of materials and properties from superconductors scientific literature
arXiv:2210.15600 · doi:10.1080/27660400.2022.2153633
Abstract
The automatic extraction of materials and related properties from the scientific literature is gaining attention in data-driven materials science (Materials Informatics). In this paper, we discuss Grobid-superconductors, our solution for automatically extracting superconductor material names and respective properties from text. Built as a Grobid module, it combines machine learning and heuristic approaches in a multi-step architecture that supports input data as raw text or PDF documents. Using Grobid-superconductors, we built SuperCon2, a database of 40324 materials and properties records from 37700 papers. The material (or sample) information is represented by name, chemical formula, and material class, and is characterized by shape, doping, substitution variables for components, and substrate as adjoined information. The properties include the Tc superconducting critical temperature and, when available, applied pressure with the Tc measurement method.
20 pages, 11 figures, 8 tables
References in corpus (4)
- Machine Learning Guided Discovery of Gigantic Magnetocaloric Effect in HoB Near Hydrogen Liquefaction Temperature
- Frame-Semantic Parsing with Softmax-Margin Segmental RNNs and a Syntactic Scaffold
- Optimizing accuracy and efficacy in data-driven materials discovery for the solar production of hydrogen
- The Diminishing Returns of Masked Language Models to Science
Cited by in corpus (5)
- Mining experimental data from Materials Science literature with Large Language Models: an evaluation study
- Accelerating superconductor discovery through tempered deep learning of the electron-phonon spectral function
- Semi-automatic staging area for high-quality structured data extraction from scientific literature
- Superconductor discovery in the emerging paradigm of Materials Informatics
- Developing a Complete AI-Accelerated Workflow for Superconductor Discovery