activity
20242026
collaborators
Showing 2025Show all

7 papers · 1 filter

cs.CL2025

The Geometry of Creative Variability: How Credal Sets Expose Calibration Gaps in Language Models

Esteban Garces Arias, Julian Rodemann, Christian Heumann

Understanding uncertainty in large language models remains a fundamental challenge, particularly in creative tasks where multiple valid outputs exist. We present a geometric framew…

cs.CL2025

GUARD: Glocal Uncertainty-Aware Robust Decoding for Effective and Efficient Open-Ended Text Generation

Yuanhao Ding, Esteban Garces Arias, Meimingwei Li +6

Open-ended text generation faces a critical challenge: balancing coherence with diversity in LLM outputs. While contrastive search-based decoding strategies have emerged to address…

cs.CL2025

Statistical Multicriteria Evaluation of LLM-Generated Text

Esteban Garces Arias, Hannah Blocher, Julian Rodemann +2

Assessing the quality of LLM-generated text remains a fundamental challenge in natural language processing. Current evaluation approaches often rely on isolated metrics or simplist…

cs.CL2025

Unveiling Factors for Enhanced POS Tagging: A Study of Low-Resource Medieval Romance Languages

Matthias Schöffel, Esteban Garces Arias, Marinus Wiedner +4

Part-of-speech (POS) tagging remains a foundational component in natural language processing pipelines, particularly critical for historical text analysis at the intersection of co…

cs.CL2025

Towards Better Open-Ended Text Generation: A Multicriteria Evaluation Framework

Esteban Garces Arias, Hannah Blocher, Julian Rodemann +3

Open-ended text generation has become a prominent task in natural language processing due to the rise of powerful (large) language models. However, evaluating the quality of these…

cs.AI2025

A Statistical Case Against Empirical Human-AI Alignment

Julian Rodemann, Esteban Garces Arias, Christoph Luther +2

Empirical human-AI alignment aims to make AI systems act in line with observed human behavior. While noble in its goals, we argue that empirical alignment can inadvertently introdu…