12 papers
Overview of the PsyDefDetect Shared Task at BioNLP 2026: Detecting Levels of Psychological Defense Mechanisms in Supportive Conversations
Hongbin Na, Zimu Wang, Zhaoming Chen +8
We present an overview of PsyDefDetect, the shared task on detecting levels of psychological defense mechanisms in emotional support dialogues, co-located with BioNLP@ACL 2026. Gro…
SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos
Xiyang Huang, Jiawei Lin, Keying Wu +9
Current video benchmarks for multimodal large language models (MLLMs) focus on event recognition, temporal ordering, and long-context recall, but overlook a harder capability requi…
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
Yunna Cai, Fan Wang, Haowei Wang +5
Evaluating the safety alignment of LLM responses in high-risk mental health dialogues is particularly difficult due to missing gold-standard answers and the ethically sensitive nat…
MiraMind: Benchmarking Reliable Mental Health Reasoning beyond Answer Accuracy
Mengxi Xiao, Kailai Yang, Pengde Zhao +12
Mental-health reasoning with large language models (LLMs) is an evidence-constrained judgment problem: models must transform limited, subjective, and often ambiguous evidence into…
Natural Language Processing for Cardiology: A Narrative Review
Kailai Yang, Yan Leng, Xin Zhang +5
Cardiovascular diseases are becoming increasingly prevalent in modern society, with a profound impact on global health and well-being. These Cardiovascular disorders are complex an…
Rumor Detection by Multi-task Suffix Learning based on Time-series Dual Sentiments
Zhiwei Liu, Kailai Yang, Eduard Hovy +1
The widespread dissemination of rumors on social media has a significant impact on people's lives, potentially leading to public panic and fear. Rumors often evoke specific sentime…