12 papers
LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap
Yining Hua, Cyrus Ayubcha, Hongbin Na +4
Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimiz…
CARE-Bench: Benchmarking Patient-Facing LLM Triage
Yining Hua, Hongbin Na, Cyrus Ayubcha
Patient-facing medical LLMs and agents increasingly answer symptom questions before clinician contact, where the key safety question is what action the user should take next. We in…
Overview of the PsyDefDetect Shared Task at BioNLP 2026: Detecting Levels of Psychological Defense Mechanisms in Supportive Conversations
Hongbin Na, Zimu Wang, Zhaoming Chen +8
We present an overview of PsyDefDetect, the shared task on detecting levels of psychological defense mechanisms in emotional support dialogues, co-located with BioNLP@ACL 2026. Gro…
Design and Report Benchmarks for Knowledge Work
Yining Hua, Hongbin Na, Cyrus Ayubcha +1
The development of LLM agents has led to a growing body of work on knowledge-work AI, including coding, research, and healthcare. However, current knowledge-work evaluation and ben…
Multi-Session Client-Centered Treatment Outcome Evaluation in Psychotherapy
Hongbin Na, Tao Shen, Shumao Yu +1
In psychotherapy, therapeutic outcome assessment, or treatment outcome evaluation, is essential to mental health care by systematically evaluating therapeutic processes and outcome…
You Never Know a Person, You Only Know Their Defenses: Detecting Levels of Psychological Defense Mechanisms in Supportive Conversations
Hongbin Na, Zimu Wang, Zhaoming Chen +9
Psychological defenses are strategies, often automatic, that people use to manage distress. Rigid or overuse of defenses is negatively linked to mental health and shapes what speak…