activity
20242026
collaborators

11 papers

cs.IR2026

AMAQA: A Metadata-based QA Dataset for RAG Systems

Davide Bruni, Marco Avvenuti, Nicola Tonellotto +1

Retrieval-augmented generation (RAG) systems are widely used in question-answering (QA) tasks, but current benchmarks lack metadata integration, limiting their evaluation in scenar…

cs.CY2026

Assessing the Reliability of Persona-Conditioned LLMs as Synthetic Survey Respondents

Erika Elizabeth Taday Morocho, Lorenzo Cima, Tiziano Fagni +2

Using persona-conditioned LLMs as synthetic survey respondents has become a common practice in computational social science and agent-based simulations. Yet, it remains unclear whe…

cs.CY2026

Beyond Trial-and-Error: Predicting User Abandonment After a Moderation Intervention

Benedetta Tessa, Lorenzo Cima, Amaury Trujillo +2

Current content moderation follows a reactive, trial-and-error approach, where interventions are applied and their effects are only measured post-hoc. In contrast, we introduce a p…

cs.CR2025

JPEGs Just Got Snipped: Croppable Signatures Against Deepfake Images

Pericle Perazzo, Massimiliano Mattei, Giuseppe Anastasi +4

Deepfakes are a type of synthetic media created using artificial intelligence, specifically deep learning algorithms. This technology can for example superimpose faces and voices o…

cs.CY2025

Investigating the heterogenous effects of a massive content moderation intervention via Difference-in-Differences

Lorenzo Cima, Benedetta Tessa, Stefano Cresci +2

In today's online environments, users encounter harm and abuse on a daily basis. Therefore, content moderation is crucial to ensure their safety and well-being. However, the effect…

cs.CL2025

Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets

Tommaso Giorgi, Lorenzo Cima, Tiziano Fagni +2

The rise of online platforms exacerbated the spread of hate speech, demanding scalable and effective detection. However, the accuracy of hate speech detection systems heavily relie…