Towards Interpretable Mental Health Analysis with Large Language Models
arXiv:2304.03347 · doi:10.18653/v1/2023.emnlp-main.370
Abstract
The latest large language models (LLMs) such as ChatGPT, exhibit strong capabilities in automated mental health analysis. However, existing relevant studies bear several limitations, including inadequate evaluations, lack of prompting strategies, and ignorance of exploring LLMs for explainability. To bridge these gaps, we comprehensively evaluate the mental health analysis and emotional reasoning ability of LLMs on 11 datasets across 5 tasks. We explore the effects of different prompting strategies with unsupervised and distantly supervised emotional information. Based on these prompts, we explore LLMs for interpretable mental health analysis by instructing them to generate explanations for each of their decisions. We convey strict human evaluations to assess the quality of the generated explanations, leading to a novel dataset with 163 human-assessed explanations. We benchmark existing automatic evaluation metrics on this dataset to guide future related works. According to the results, ChatGPT shows strong in-context learning ability but still has a significant gap with advanced task-specific methods. Careful prompt engineering with emotional cues and expert-written few-shot examples can also effectively improve performance on mental health analysis. In addition, ChatGPT generates explanations that approach human performance, showing its great potential in explainable mental health analysis.
Accepted by EMNLP 2023 main conference as a long paper
References in corpus (16)
- Training language models to follow instructions with human feedback
- LLaMA: Open and Efficient Foundation Language Models
- ChatGPT: Jack of all trades, master of none
- DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset
- BARTScore: Evaluating Generated Text as Text Generation
- A Comprehensive Capability Analysis of GPT-3 and GPT-3.5 Series Models
- Emotion Detection on TV Show Transcripts with Sequence-based Convolutional Neural Networks
- Emotion fusion for mental illness detection from social media: A survey
- Cluster-Level Contrastive Learning for Emotion Recognition in Conversations
- GPTScore: Evaluate as You Desire
- Evaluation of ChatGPT for NLP-based Mental Health Applications
- ChatGPT as a Factual Inconsistency Evaluator for Text Summarization
- Will Affective Computing Emerge from Foundation Models and General AI? A First Evaluation on ChatGPT
- How Robust is GPT-3.5 to Predecessors? A Comprehensive Study on Language Understanding Tasks
- Hierarchical Attention Network for Explainable Depression Detection on Twitter Aided by Metaphor Concept Mappings
- Psychiatric Scale Guided Risky Post Screening for Early Detection of Depression
Cited by in corpus (8)
- The opportunities and risks of large language models in mental health
- Designing Interpretable ML System to Enhance Trust in Healthcare: A Systematic Review to Proposed Responsible Clinician-AI-Collaboration Framework
- MentaLLaMA: Interpretable Mental Health Analysis on Social Media with Large Language Models
- Prompt engineering paradigms for medical applications: scoping review and recommendations for better practices
- Socially Aware Synthetic Data Generation for Suicidal Ideation Detection Using Large Language Models
- Clinical Insights: A Comprehensive Review of Language Models in Medicine
- Large Language Model for Mental Health: A Systematic Review
- Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges