2 papers
cs.IR2026
SciDef: Datasets and Tools for Automated Definition Extraction from Scientific Literature with LLMs
Filip KuÄera, Christoph Mandl, Isao Echizen +2
Scientific concepts are often defined inconsistently across papers, making it difficult to compare findings, reuse terminology, and build reliable downstream resources. We present…
cs.CL2025
The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection
Tomas Horych, Christoph Mandl, Terry Ruas +4
High annotation costs from hiring or crowdsourcing complicate the creation of large, high-quality datasets needed for training reliable text classifiers. Recent research suggests u…