activity
20242026
collaborators

8 papers

cs.CL2026

Beyond "AI Language": The case for the idiolectal nature of LLM output

Karolina Rudnicka, Thomas Stephan Juzek

While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that this perspective coexists with distinct, mod…

cs.CL2026

Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models

Thomas Stephan Juzek, Xiaoyang Ming, Jose A. Hernandez

The language used by digital chat assistants such as ChatGPT can diverge from human expectations (misalignment). Research, mostly on Scientific English, has described both what div…

cs.CL2026

Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning

Xiaoyang Ming, Jose Hernandez, Thomas Stephan Juzek

Various language domains have undergone remarkable changes in recent years; these shifts are largely attributed to the advent of Large Language Models and their misalignment with n…

cs.CL2026

AI-Associated Lexical Shifts Across 34 Languages: Cross-Lingual Convergence and Diachronic Uptake in News Writing

Thomas Stephan Juzek

AI-associated lexical shifts have been documented mainly in Scientific English. We extend this work to 34 languages in the WMT News Crawl corpus, refining a split-halves continuati…

cs.CL2025

Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback

Tom S. Juzek, Zina B. Ward

Large Language Models (LLMs) are known to overuse certain terms like "delve" and "intricate." The exact reasons for these lexical choices, however, have been unclear. Using Meta's…

cs.CL2025

Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English

Bryce Anderson, Riley Galpin, Tom S. Juzek

In recent years, written language, particularly in science and education, has undergone remarkable shifts in word usage. These changes are widely attributed to the growing influenc…