15 papers
When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge
Xin Sun, Di Wu, Yuchen Guo +4
LLM-as-Judge systems can produce multi-dimensional evaluations, such as trustworthiness, reliability, and factuality, and these outputs are often interpreted as independent evidenc…
MedEasy: Designing AI Standardized Patients for Clinical Consultation Training
Zhiqi Gao, Huarui Luo, Guo Zhu +6
AI standardized patients are becoming a setting for professional training in clinical consultation. This paper presents MedEasy, a multi-agent system that organizes virtual-patient…
Trust Stack for Mental Health AI: A Survey of Calibration across Human, Interaction, and AI Layers
Xin Sun, Yue Su, Yifan Mo +9
Language-based AI is increasingly deployed for mental health support, yet trust is evaluated in interdisciplinary but operationally misaligned ways: NLP and AI work measures robust…
SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity
Yifan Mo, Xiao Fu, Yue Su +4
This work investigates the ability of large language models (LLMs) to generate mathematical equations from scientific texts. Prior work faces challenges in unstructured grounding,…
Herculean: An Agentic Benchmark for Financial Intelligence
Xueqing Peng, Zhuohan Xie, Yupeng Cao +60
As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…
StoryMI: Steerable Multi-Agent Therapeutic Dialogue Generation
Qingyu Meng, Min Chen, Dingming Liu +5
Large language models (LLMs) can generate fluent dialogue, but prior works lack situational grounding, dynamic strategy control, and evaluation aligned with clinical standards in m…