Publications (38)
Entity Linking using LLMs for Automated Product Carbon Footprint Estimation
Steffen Castle, Julian Moreno Schneider, Leonhard Hennig +1
Growing concerns about climate change and sustainability are driving manufacturers to take significant steps toward reducing their carbon footprints. For these manufacturers, a fir…
Toward FAIR Semantic Publishing of Research Dataset Metadata in the Open Research Knowledge Graph
Raia Abu Ahmad, Jennifer D'Souza, Matthäus Zloch +5
Search engines these days can serve datasets as search results. Datasets get picked up by search technologies based on structured descriptions on their official web pages, informed…
A Dataset of German Legal Documents for Named Entity Recognition
Elena Leitner, Georg Rehm, Julián Moreno-Schneider
We describe a dataset developed for Named Entity Recognition in German federal court decisions. It consists of approx. 67,000 sentences with over 2 million tokens. The resource con…
SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing
Luca Foppiano, Sotaro Takeshita, Pedro Ortiz Suarez +6
SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English…
Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data
Ekaterina Borisova, Fabio Barth, Nils Feldhus +5
Tables are among the most widely used tools for representing structured data in research, business, medicine, and education. Although LLMs demonstrate strong performance in downstr…
Evaluating Document Representations for Content-based Legal Literature Recommendations
Malte Ostendorff, Elliott Ash, Terry Ruas +3
Recommender systems assist legal professionals in finding relevant literature for supporting their case. Despite its importance for the profession, legal applications do not reflec…