most citedSemEval-2023 Task 10: Explainable Detection of Online Sexism

48 citations · 80 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2024

Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models

Fabio Pernisi, Dirk Hovy, Paul Röttger

As diverse linguistic communities and users adopt large language models (LLMs), assessing their safety across languages becomes critical. Despite ongoing efforts to make LLMs safe,…

cs.CL20243 cited

Evidence of a log scaling law for political persuasion with large language models

Kobi Hackenburg, Ben M. Tappin, Paul Röttger +3

Large language models can now generate political messages as persuasive as those written by humans, raising concerns about how far this persuasiveness may continue to increase with…

cs.CL20242 cited

Beyond Flesch-Kincaid: Prompt-based Metrics Improve Difficulty Classification of Educational Texts

Donya Rooein, Paul Rottger, Anastassia Shaitarova +1

Using large language models (LLMs) for educational applications like dialogue-based teaching is a hot topic. Effective teaching, however, requires teachers to adapt the difficulty…

cs.CL2024

Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech Dataset

Janis Goldzycher, Paul Röttger, Gerold Schneider

Hate speech detection models are only as good as the data they are trained on. Datasets sourced from social media suffer from systematic gaps and biases, leading to unreliable mode…

cs.CL20233 cited

The Empty Signifier Problem: Towards Clearer Paradigms for Operationalising "Alignment" in Large Language Models

Hannah Rose Kirk, Bertie Vidgen, Paul Röttger +1

In this paper, we address the concept of "alignment" in large language models (LLMs) through the lens of post-structuralist socio-political theory, specifically examining its paral…

cs.CL2023

The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values

Hannah Rose Kirk, Andrew M. Bean, Bertie Vidgen +2

Human feedback is increasingly used to steer the behaviours of Large Language Models (LLMs). However, it is unclear how to collect and incorporate feedback in a way that is efficie…