collaborators

6 papers

cs.CL2026

DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots

Jared Moore, Andrea Mock, Yifan Mai +9

Mental health professionals have raised concerns about risks of psychological harm from interaction with large language models (LLMs), including "delusional spirals" in which conce…

cs.CL2026

Characterizing Delusional Spirals through Human-LLM Chat Logs

Jared Moore, Ashish Mehta, William Agnew +11

As large language models (LLMs) have proliferated, disturbing anecdotal reports of negative psychological effects, such as delusions, self-harm, and ``AI psychosis,'' have emerged…

cs.HC2026

Can LLM-Simulated Practice and Feedback Upskill Human Counselors? A Randomized Study with 90+ Novice Counselors

Ryan Louie, Raj Sanjay Shah, Ifdita Hasan Orney +3

The growing demand for accessible mental health support requires training more counselors, yet existing approaches remain resource-intensive and difficult to scale. LLMs can realis…

cs.CL2026

TherapyGym: Evaluating and Aligning Clinical Fidelity and Safety in Therapy Chatbots

Fangrui Huang, Souhad Chbeir, Arpandeep Khatua +8

Large language models (LLMs) are increasingly used for mental-health support; yet prevailing evaluation methods--fluency metrics, preference tests, and generic dialogue benchmarks-…

cs.HC2025

Conversational Self-Play for Discovering and Understanding Psychotherapy Approaches

Onno P Kampman, Michael Xing, Charmaine Lim +4

This paper explores conversational self-play with LLMs as a scalable approach for analyzing and exploring psychotherapy approaches, evaluating how well AI-generated therapeutic dia…

cs.HC2025

SPHERE: An Evaluation Card for Human-AI Systems

Qianou Ma, Dora Zhao, Xinran Zhao +6

In the era of Large Language Models (LLMs), establishing effective evaluation methods and standards for diverse human-AI interaction systems is increasingly challenging. To encoura…