activity
20242026
collaborators

8 papers

cs.AI2026

LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap

Yining Hua, Cyrus Ayubcha, Hongbin Na +4

Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimiz…

cs.AI2026

CARE-Bench: Benchmarking Patient-Facing LLM Triage

Yining Hua, Hongbin Na, Cyrus Ayubcha

Patient-facing medical LLMs and agents increasingly answer symptom questions before clinician contact, where the key safety question is what action the user should take next. We in…

cs.CL2026

Overview of the PsyDefDetect Shared Task at BioNLP 2026: Detecting Levels of Psychological Defense Mechanisms in Supportive Conversations

Hongbin Na, Zimu Wang, Zhaoming Chen +8

We present an overview of PsyDefDetect, the shared task on detecting levels of psychological defense mechanisms in emotional support dialogues, co-located with BioNLP@ACL 2026. Gro…

cs.AI2026

Designing Benchmarks for Knowledge Work

Yining Hua, Hongbin Na, Cyrus Ayubcha +1

AI agents are moving quickly from answering isolated questions toward completing work through tools, software environments, and multi-step workflows. Much of what these systems are…

cs.CL2025

You Never Know a Person, You Only Know Their Defenses: Detecting Levels of Psychological Defense Mechanisms in Supportive Conversations

Hongbin Na, Zimu Wang, Zhaoming Chen +9

Psychological defenses are strategies, often automatic, that people use to manage distress. Rigid or overuse of defenses is negatively linked to mental health and shapes what speak…

cs.HC2025

MindBenchAI: An Actionable Platform to Evaluate the Profile and Performance of Large Language Models in a Mental Healthcare Context

Bridget Dwyer, Matthew Flathers, Akane Sano +28

Individuals are increasingly utilizing large language model (LLM)based tools for mental health guidance and crisis support in place of human experts. While AI technology has great…