17 papers
Breaking the Likelihood Trap: Variance-Calibrated Modulation for Large Language Model Decoding
Yuanhao Ding, Meimingwei Li, Esteban Garces Arias +3
In open-ended generation, LLMs frequently fall into the "likelihood trap", marked by repetitive degeneration and vocabulary dullness, creating a discrepancy between machine-generat…
The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices
Esteban Garces Arias, Nurzhan Sapargali, Christian Heumann +1
Why does machine-generated text remain detectable? We trace the answer to the decoding stage: standard strategies such as top- and nucleus sampling restrict generation to high-p…
Automated sign detection across the Electronic Babylonian Library: A large-scale dataset and end-to-end cuneiform OCR pipeline
Wentao Che, Esteban Garcés Arias, Asim Niaz +2
Learning to read cuneiform tablets is an extremely demanding task; consequently, of the roughly half million excavated tablets, only a small fraction has been analysed by Assyriolo…
Lost in Translation? Exploring the Shift in Grammatical Gender from Latin to Occitan
Ahan Chatterjee, Matthias Schöffel, Matthias AÃenmacher +2
The diachronic evolution from Latin to the Romance languages involved a restructuring of the grammatical gender system from a tripartite configuration (masculine, feminine, neuter)…
Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansion
Meimingwei Li, Yuanhao Ding, Esteban Garces Arias +1
Recent work has identified a counterintuitive phenomenon termed "Hyperfitting", where fine-tuning Large Language Models (LLMs) to near-zero training loss on small datasets surprisi…
From Traditional Taggers to LLMs: A Comparative Study of POS Tagging for Medieval Romance Languages
Matthias Schöffel, Esteban Garces Arias
Part-of-speech (POS) tagging for Medieval Romance languages remains challenging due to orthographic variation, morphological complexity, and limited annotated resources. This paper…