4 papers
An Extreme Multi-label Text Classification (XMTC) Library Dataset: What if we took "Use of Practical AI in Digital Libraries" seriously?
Jennifer D'Souza, Sameer Sadruddin, Maximilian Kähler +5
Subject indexing is vital for discovery but hard to sustain at scale and across languages. We release a large bilingual (English/German) corpus of catalog records annotated with th…
Diagnosing Structural Failures in LLM-Based Evidence Extraction for Meta-Analysis
Zhiyin Tan, Jennifer D'Souza
Systematic reviews and meta-analyses rely on converting narrative articles into structured, numerically grounded study records. Despite rapid advances in large language models (LLM…
Toward Purpose-oriented Topic Model Evaluation enabled by Large Language Models
Zhiyin Tan, Jennifer D'Souza
This study presents a framework for automated evaluation of dynamically evolving topic models using Large Language Models (LLMs). Topic modeling is essential for organizing and ret…
Bridging the Evaluation Gap: Leveraging Large Language Models for Topic Model Evaluation
Zhiyin Tan, Jennifer D'Souza
This study presents a framework for automated evaluation of dynamically evolving topic taxonomies in scientific literature using Large Language Models (LLMs). In digital library sy…