3 papers
cs.CL2026
An Extreme Multi-label Text Classification (XMTC) Library Dataset: What if we took "Use of Practical AI in Digital Libraries" seriously?
Jennifer D'Souza, Sameer Sadruddin, Maximilian Kähler +5
Subject indexing is vital for discovery but hard to sustain at scale and across languages. We release a large bilingual (English/German) corpus of catalog records annotated with th…
cs.CL2025
Annif at the GermEval-2025 LLMs4Subjects Task: Traditional XMTC Augmented by Efficient LLMs
Osma Suominen, Juho Inkinen, Mona Lehtinen
This paper presents the Annif system in the LLMs4Subjects shared task (Subtask 2) at GermEval-2025. The task required creating subject predictions for bibliographic records using l…
cs.CL2025
Annif at SemEval-2025 Task 5: Traditional XMTC augmented by LLMs
Osma Suominen, Juho Inkinen, Mona Lehtinen
This paper presents the Annif system in SemEval-2025 Task 5 (LLMs4Subjects), which focussed on subject indexing using large language models (LLMs). The task required creating subje…