Evaluation of ChatGPT for NLP-based Mental Health Applications
arXiv:2303.15727
Abstract
Large language models (LLM) have been successful in several natural language understanding tasks and could be relevant for natural language processing (NLP)-based mental health application research. In this work, we report the performance of LLM-based ChatGPT (with gpt-3.5-turbo backend) in three text-based mental health classification tasks: stress detection (2-class classification), depression detection (2-class classification), and suicidality detection (5-class classification). We obtained annotated social media posts for the three classification tasks from public datasets. Then ChatGPT API classified the social media posts with an input prompt for classification. We obtained F1 scores of 0.73, 0.86, and 0.37 for stress detection, depression detection, and suicidality detection, respectively. A baseline model that always predicted the dominant class resulted in F1 scores of 0.35, 0.60, and 0.19. The zero-shot classification accuracy obtained with ChatGPT indicates a potential use of language models for mental health classification tasks.
Cited by in corpus (6)
- The opportunities and risks of large language models in mental health
- Towards Interpretable Mental Health Analysis with Large Language Models
- Exploring the Frontiers of LLMs in Psychological Applications: A Comprehensive Review
- PromptMTopic: Unsupervised Multimodal Topic Modeling of Memes using Large Language Models
- Survey and Experiments on Mental Disorder Detection via Social Media: From Large Language Models and RAG to Agents
- Conceptualizing Suicidal Behavior: Utilizing Explanations of Predicted Outcomes to Analyze Longitudinal Social Media Data