From the 1 of 5 linked papers with an AI index.
5 papers
SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions
Jennifer D'Souza, Sameer Sadruddin, Anisa Rula +23
The paper introduces SciSchema.org, a multidisciplinary collection of 16 expert‑annotated schemas for describing scientific processes in a structured way, created using a human‑in‑…
An Extreme Multi-label Text Classification (XMTC) Library Dataset: What if we took "Use of Practical AI in Digital Libraries" seriously?
Jennifer D'Souza, Sameer Sadruddin, Maximilian Kähler +5
Subject indexing is vital for discovery but hard to sustain at scale and across languages. We release a large bilingual (English/German) corpus of catalog records annotated with th…
NFDI4DS Shared Tasks for Scholarly Document Processing
Raia Abu Ahmad, Rana Abdulla, Tilahun Abedissa Taffa +18
Shared tasks are powerful tools for advancing research through community-based standardised evaluation. As such, they play a key role in promoting findable, accessible, interoperab…
SemEval-2025 Task 5: LLMs4Subjects -- LLM-based Automated Subject Tagging for a National Technical Library's Open-Access Catalog
Jennifer D'Souza, Sameer Sadruddin, Holger Israel +2
We present SemEval-2025 Task 5: LLMs4Subjects, a shared task on automated subject tagging for scientific and technical records in English and German using the GND taxonomy. Partici…
LLMs4SchemaDiscovery: A Human-in-the-Loop Workflow for Scientific Schema Mining with Large Language Models
Sameer Sadruddin, Jennifer D'Souza, Eleni Poupaki +7
Extracting structured information from unstructured text is crucial for modeling real-world processes, but traditional schema mining relies on semi-structured data, limiting scalab…