3 papers
cs.LG2026
Improving Certified Robustness via Adversarial Distillation
Matteo Melis, Jesus Martinez Del Rincon, Vishal Sharma
Certified training aims to produce models whose predictions can be formally verified against adversarial perturbations, typically by optimising upper bounds on the worst-case loss…
cs.CL2025
A Modular Taxonomy for Hate Speech Definitions and Its Impact on Zero-Shot LLM Classification Performance
Matteo Melis, Gabriella Lapesa, Dennis Assenmacher
Detecting harmful content is a crucial task in the landscape of NLP applications for Social Good, with hate speech being one of its most dangerous forms. But what do we mean by hat…
cs.CL2025
Tell Me What You Know About Sexism: Expert-LLM Interaction Strategies and Co-Created Definitions for Zero-Shot Sexism Detection
Myrthe Reuver, Indira Sen, Matteo Melis +1
This paper investigates hybrid intelligence and collaboration between researchers of sexism and Large Language Models (LLMs), with a four-component pipeline. First, nine sexism res…