3 papers
cs.CV2026
Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking
Timothee Mickus, Claudio Savelli, Eduardo Calò +10
In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-gene…
cs.SE2026
Evaluating LLM-Based Goal Extraction in Requirements Engineering: Prompting Strategies and Their Limitations
Anna Arnaudo, Riccardo Coppola, Maurizio Morisio +4
Due to the textual and repetitive nature of many Requirements Engineering (RE) artefacts, Large Language Models (LLMs) have proven useful to automate their generation and processin…
cs.AI2026
Analysis Of Linguistic Stereotypes in Single and Multi-Agent Generative AI Architectures
Martina Ullasci, Marco Rondina, Riccardo Coppola +6
Many works in the literature show that LLM outputs exhibit discriminatory behaviour, triggering stereotype-based inferences based on the dialect in which the inputs are written. Th…