3 papers
cs.CL2026
HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
Matthew Flathers, Phuong Anh Nguyen, Jill Noorily +7
General-purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are not always resolved by clinical specialty, making domain-specific performan…
cs.CY2026
Depictions of Depression in Generative AI Video Models: A Preliminary Study of OpenAI's Sora 2
Matthew Flathers, Griffin Smith, Julian Herpertz +2
Generative video models are increasingly capable of producing complex depictions of mental health experiences, yet little is known about how these systems represent conditions like…
cs.HC2025
MindBenchAI: An Actionable Platform to Evaluate the Profile and Performance of Large Language Models in a Mental Healthcare Context
Bridget Dwyer, Matthew Flathers, Akane Sano +28
Individuals are increasingly utilizing large language model (LLM)based tools for mental health guidance and crisis support in place of human experts. While AI technology has great…