collaborators

15 papers

cs.CL2026

Breaking the Likelihood Trap: Variance-Calibrated Modulation for Large Language Model Decoding

Yuanhao Ding, Meimingwei Li, Esteban Garces Arias +3

In open-ended generation, LLMs frequently fall into the "likelihood trap", marked by repetitive degeneration and vocabulary dullness, creating a discrepancy between machine-generat…

cs.CL2026

The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices

Esteban Garces Arias, Nurzhan Sapargali, Christian Heumann +1

Why does machine-generated text remain detectable? We trace the answer to the decoding stage: standard strategies such as top- and nucleus sampling restrict generation to high-p…

cs.CL2026

Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansion

Meimingwei Li, Yuanhao Ding, Esteban Garces Arias +1

Recent work has identified a counterintuitive phenomenon termed "Hyperfitting", where fine-tuning Large Language Models (LLMs) to near-zero training loss on small datasets surprisi…

cs.LG2026

Self-Reinforcing Controllable Synthesis of Rare Relational Data via Bayesian Calibration

Chongsheng Zhang, Hao Wang, Zelong Yu +7

Imbalanced data are commonly present in real-world applications. While data synthesis can effectively mitigate data scarcity for rare classes, and LLMs have revolutionized text gen…

cs.AI2026

Min- Sampling: Decoupling Truncation from Temperature Scaling via Relative Logit Dynamics

Yuanhao Ding, Meimingwei Li, Esteban Garces Arias +3

The quality of text generated by large language models depends critically on the decoding sampling strategy. While mainstream methods such as Top-, Top-, and Min- achieve…

cs.CL2025

Fair Play in the Newsroom: Actor-Based Filtering Gender Discrimination in Text Corpora

Stefanie Urchs, Veronika Thurner, Matthias Aßenmacher +2

Language corpora are the foundation of most natural language processing research, yet they often reproduce structural inequalities. One such inequality is gender discrimination in…