From Classification to Clinical Insights: Towards Analyzing and Reasoning About Mobile and Behavioral Health Data With Large Language Models
arXiv:2311.13063 · doi:10.1145/3659604
Abstract
Passively collected behavioral health data from ubiquitous sensors holds significant promise to provide mental health professionals insights from patient's daily lives; however, developing analysis tools to use this data in clinical practice requires addressing challenges of generalization across devices and weak or ambiguous correlations between the measured signals and an individual's mental health. To address these challenges, we take a novel approach that leverages large language models (LLMs) to synthesize clinically useful insights from multi-sensor data. We develop chain of thought prompting methods that use LLMs to generate reasoning about how trends in data such as step count and sleep relate to conditions like depression and anxiety. We first demonstrate binary depression classification with LLMs achieving accuracies of 61.1% which exceed the state of the art. While it is not robust for clinical use, this leads us to our key finding: even more impactful and valued than classification is a new human-AI collaboration approach in which clinician experts interactively query these tools and combine their domain expertise and context about the patient with AI generated reasoning to support clinical decision-making. We find models like GPT-4 correctly reference numerical data 75% of the time, and clinician participants express strong interest in using this approach to interpret self-tracking data.
References in corpus (7)
- ChatGPT: Jack of all trades, master of none
- Lost in the Middle: How Language Models Use Long Contexts
- Llemma: An Open Language Model For Mathematics
- LLM Augmented LLMs: Expanding Capabilities through Composition
- Probing Explicit and Implicit Gender Bias through LLM Conditional Text Generation
- On Second Thought, Let's Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning
- BBQ: A Hand-Built Bias Benchmark for Question Answering
Cited by in corpus (8)
- Large Language Models for Wearable Sensor-Based Human Activity Recognition, Health Monitoring, and Behavioral Modeling: A Survey of Early Trends, Datasets, and Challenges
- Artificial Intelligence of Things: A Survey
- GPTCoach: Towards LLM-Based Physical Activity Coaching
- Exploring Personalized Health Support through Data-Driven, Theory-Guided LLMs: A Case Study in Sleep Health
- Vision-Based Multimodal Interfaces: A Survey and Taxonomy for Enhanced Context-Aware System Design
- Evaluating Large Language Models as Virtual Annotators for Time-series Physical Sensing Data
- Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-In-The-Loop LLM
- Deep Learning-Based Detection of Cognitive Impairment from Passive Smartphone Sensing with Routine-Aware Augmentation and Demographic Personalization