12 papers
Fair Play in the Newsroom: Actor-Based Filtering Gender Discrimination in Text Corpora
Stefanie Urchs, Veronika Thurner, Matthias AÃenmacher +2
Language corpora are the foundation of most natural language processing research, yet they often reproduce structural inequalities. One such inequality is gender discrimination in…
GUARD: Glocal Uncertainty-Aware Robust Decoding for Effective and Efficient Open-Ended Text Generation
Yuanhao Ding, Esteban Garces Arias, Meimingwei Li +6
Open-ended text generation faces a critical challenge: balancing coherence with diversity in LLM outputs. While contrastive search-based decoding strategies have emerged to address…
Are All Genders Equal in the Eyes of Algorithms? -- Analysing Search and Retrieval Algorithms for Algorithmic Gender Fairness
Stefanie Urchs, Veronika Thurner, Matthias AÃenmacher +3
Algorithmic systems such as search engines and information retrieval platforms significantly influence academic visibility and the dissemination of knowledge. Despite assumptions o…
Explainable Coarse-to-Fine Ancient Manuscript Duplicates Discovery
Chongsheng Zhang, Shuwen Wu, Yingqi Chen +5
Ancient manuscripts are the primary source of ancient linguistic corpora. However, many ancient manuscripts exhibit duplications due to unintentional repeated publication or delibe…
Statistical Multicriteria Evaluation of LLM-Generated Text
Esteban Garces Arias, Hannah Blocher, Julian Rodemann +2
Assessing the quality of LLM-generated text remains a fundamental challenge in natural language processing. Current evaluation approaches often rely on isolated metrics or simplist…
Unveiling Factors for Enhanced POS Tagging: A Study of Low-Resource Medieval Romance Languages
Matthias Schöffel, Esteban Garces Arias, Marinus Wiedner +4
Part-of-speech (POS) tagging remains a foundational component in natural language processing pipelines, particularly critical for historical text analysis at the intersection of co…