6 papers · 1 filter
TailNLG: A Multilingual Benchmark Addressing Verbalization of Long-Tail Entities
Lia Draetta, Michael Oliverio, Virginia Ramón-Ferrer +6
The automatic verbalization of structured knowledge is a key task for making knowledge graphs accessible to non-expert users and supporting retrieval-augmented generation systems.…
Conspiracy Frame: a Semiotically-Driven Approach for Conspiracy Theories Detection
Heidi Campana Piva, Shaina Ashraf, Maziar Kianimoghadam Jouneghani +4
Conspiracy theories are anti-authoritarian narratives that lead to social conflict, impacting how people perceive political information. To help in understanding this issue, we int…
Are you sure? Measuring models bias in content moderation through uncertainty
Alessandra Urbinati, Mirko Lai, Simona Frenda +1
Automatic content moderation is crucial to ensuring safety in social media. Language Model-based classifiers are being increasingly adopted for this task, but it has been shown tha…
What Are They Filtering Out? An Experimental Benchmark of Filtering Strategies for Harm Reduction in Pretraining Datasets
Marco Antonio Stranisci, Christian Hardmeier
Data filtering strategies are a crucial component to develop safe Large Language Models (LLM), since they support the removal of harmful contents from pretraining datasets. There i…
Dealing with Controversy: An Emotion and Coping Strategy Corpus Based on Role Playing
Enrica Troiano, Sofie Labat, Marco Antonio Stranisci +3
There is a mismatch between psychological and computational studies on emotions. Psychological research aims at explaining and documenting internal mechanisms of these phenomena, w…
APPReddit: a Corpus of Reddit Posts Annotated for Appraisal
Marco Antonio Stranisci, Simona Frenda, Eleonora Ceccaldi +3
Despite the large number of computational resources for emotion recognition, there is a lack of data sets relying on appraisal models. According to Appraisal theories, emotions are…