2 papers
cs.CL2026
ThreatCore: A Benchmark for Explicit and Implicit Threat Detection
Davide Bruni, Carlo Bardazzi, Maurizio Tesconi
Threat detection in Natural Language Processing lacks consistent definitions and standardized benchmarks, and is often conflated with broader phenomena such as toxicity, hate speec…
cs.IR2026
AMAQA: A Metadata-based QA Dataset for RAG Systems
Davide Bruni, Marco Avvenuti, Nicola Tonellotto +1
Retrieval-augmented generation (RAG) systems are widely used in question-answering (QA) tasks, but current benchmarks lack metadata integration, limiting their evaluation in scenar…