8 papers
Beyond "AI Language": The case for the idiolectal nature of LLM output
Karolina Rudnicka, Thomas Stephan Juzek
While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that this perspective coexists with distinct, mod…
Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models
Thomas Stephan Juzek, Xiaoyang Ming, Jose A. Hernandez
The language used by digital chat assistants such as ChatGPT can diverge from human expectations (misalignment). Research, mostly on Scientific English, has described both what div…
Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning
Xiaoyang Ming, Jose Hernandez, Thomas Stephan Juzek
Various language domains have undergone remarkable changes in recent years; these shifts are largely attributed to the advent of Large Language Models and their misalignment with n…
AI-Associated Lexical Shifts Across 34 Languages: Cross-Lingual Convergence and Diachronic Uptake in News Writing
Thomas Stephan Juzek
AI-associated lexical shifts have been documented mainly in Scientific English. We extend this work to 34 languages in the WMT News Crawl corpus, refining a split-halves continuati…
Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback
Tom S. Juzek, Zina B. Ward
Large Language Models (LLMs) are known to overuse certain terms like "delve" and "intricate." The exact reasons for these lexical choices, however, have been unclear. Using Meta's…
Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English
Bryce Anderson, Riley Galpin, Tom S. Juzek
In recent years, written language, particularly in science and education, has undergone remarkable shifts in word usage. These changes are widely attributed to the growing influenc…