1 paper
Pablo A. Fonseca, Raquel Rodríguez-Carvajal, Rafael A. Calvo
Large language models are increasingly consulted at moments of distress, yet single-turn benchmarks neither test sustained exchanges nor distinguish between users. We built a perso…