4 papers · 1 filter
Beyond the pale: Assessing prevalence and contents of extremist speech in LLM training data
Dmitry Nikolaev, Ashley A. Mattheis
Despite a strong interest on the part of the research community in the topic of trustworthy and safe AI, the composition of the text corpora that large language models (LLMs) encou…
What Is The Political Content in LLMs' Pre- and Post-Training Data?
Tanise Ceron, Dmitry Nikolaev, Dominik Stammbach +1
Large language models (LLMs) are known to generate politically biased text. Yet, it remains unclear how such biases arise, making it difficult to design effective mitigation strate…
Strategies for political-statement segmentation and labelling in unstructured text
Dmitry Nikolaev, Sean Papay
Analysis of parliamentary speeches and political-party manifestos has become an integral area of computational study of political texts. While speeches have been overwhelmingly ana…
Classifier identification in Ancient Egyptian as a low-resource sequence-labelling task
Dmitry Nikolaev, Jorke Grotenhuis, Haleli Harel +1
The complex Ancient Egyptian (AE) writing system was characterised by widespread use of graphemic classifiers (determinatives): silent (unpronounced) hieroglyphic signs clarifying…