1 paper
Abeer Badawi, Moyosoreoluwa Olatosi, Negin Baghbanzadeh +5
Recent incidents involving LLMs used for mental-health support reveal a critical evaluation gap: surface-level safety scores do not capture how models behave across realistic, emot…