3 papers
cs.CL2026
Beyond the pale: Assessing prevalence and contents of extremist speech in LLM training data
Dmitry Nikolaev, Ashley A. Mattheis
Despite a strong interest on the part of the research community in the topic of trustworthy and safe AI, the composition of the text corpora that large language models (LLMs) encou…
cs.CL2026
What Is The Political Content in LLMs' Pre- and Post-Training Data?
Tanise Ceron, Dmitry Nikolaev, Dominik Stammbach +1
Large language models (LLMs) are known to generate politically biased text. Yet, it remains unclear how such biases arise, making it difficult to design effective mitigation strate…
cs.CL2025
Strategies for political-statement segmentation and labelling in unstructured text
Dmitry Nikolaev, Sean Papay
Analysis of parliamentary speeches and political-party manifestos has become an integral area of computational study of political texts. While speeches have been overwhelmingly ana…