2 papers
cs.HC2026
All four leading LLMs talk more than they listen to personality-verified synthetic help-seekers
Pablo A. Fonseca, Raquel Rodríguez-Carvajal, Rafael A. Calvo
Large language models are increasingly consulted at moments of distress, yet single-turn benchmarks neither test sustained exchanges nor distinguish between users. We built a perso…
cs.HC2026
Evaluating Social Engineering Risks in AI-based Interaction using Biometrics and a Gaming Setup
Roberto Daza, Javier Irigoyen, Ivan Lopez +5
We introduce AIriskEval-gaming, an open platform and dataset to evaluate social engineering risks in LLM-mediated multimodal interaction through controlled games. It supports human…