2 papers
cs.CL2026
K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
Laura M. Vowels, Matthew J. Vowels, Shivali Sharma +9
% !TEX root = ../main.tex People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk conversations remains poorly chara…
cs.CY2026
Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs
Alireza A. Safaei, Laura M. Vowels, Matthew J. Vowels +2
The deployment of large language models (LLMs) in mental health contexts raises questions about the relationship between clinical safety and environmental cost. In this paper, we e…