annotator disagreement 1conformal prediction 1content moderation 1demographic bias 1large language models 1uncertainty estimation 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CL2026
The Ghost Annotator: a Framework to Explore Human Label Variation in Content Moderation through Conformal Prediction
Mirko Lai, Alessandra Urbinati, Simona Frenda +2
The paper proposes a framework that uses conformal prediction and collaborative‑filtering style annotator representations to study how large language models agree or disagree with…
cs.CL2025
Are you sure? Measuring models bias in content moderation through uncertainty
Alessandra Urbinati, Mirko Lai, Simona Frenda +1
Automatic content moderation is crucial to ensuring safety in social media. Language Model-based classifiers are being increasingly adopted for this task, but it has been shown tha…