3 papers
cs.CL2026
Next Reply Prediction X Dataset: Linguistic Discrepancies in Naively Generated Content
Simon Münker, Nils Schwager, Kai Kugler +2
The increasing use of Large Language Models (LLMs) as proxies for human participants in social science research presents a promising, yet methodologically risky, paradigm shift. Wh…
cs.CL2025
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
Kai Kugler
We present the first systematic investigation of Martin's Law - the empirical relationship between word frequency and polysemy - in text generated by neural language models during…
cs.CL2024
Zero-shot prompt-based classification: topic labeling in times of foundation models in German Tweets
Simon Münker, Kai Kugler, Achim Rettinger
Filtering and annotating textual data are routine tasks in many areas, like social media or news analytics. Automating these tasks allows to scale the analyses wrt. speed and bread…