activity
20242026
collaborators

15 papers

cs.AI2026

When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge

Xin Sun, Di Wu, Yuchen Guo +4

LLM-as-Judge systems can produce multi-dimensional evaluations, such as trustworthiness, reliability, and factuality, and these outputs are often interpreted as independent evidenc…

cs.HC2026

MedEasy: Designing AI Standardized Patients for Clinical Consultation Training

Zhiqi Gao, Huarui Luo, Guo Zhu +6

AI standardized patients are becoming a setting for professional training in clinical consultation. This paper presents MedEasy, a multi-agent system that organizes virtual-patient…

cs.CL2026

Trust Stack for Mental Health AI: A Survey of Calibration across Human, Interaction, and AI Layers

Xin Sun, Yue Su, Yifan Mo +9

Language-based AI is increasingly deployed for mental health support, yet trust is evaluated in interdisciplinary but operationally misaligned ways: NLP and AI work measures robust…

cs.AI2026

SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity

Yifan Mo, Xiao Fu, Yue Su +4

This work investigates the ability of large language models (LLMs) to generate mathematical equations from scientific texts. Prior work faces challenges in unstructured grounding,…

cs.AI2026

Herculean: An Agentic Benchmark for Financial Intelligence

Xueqing Peng, Zhuohan Xie, Yupeng Cao +60

As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…

cs.CL2026

StoryMI: Steerable Multi-Agent Therapeutic Dialogue Generation

Qingyu Meng, Min Chen, Dingming Liu +5

Large language models (LLMs) can generate fluent dialogue, but prior works lack situational grounding, dynamic strategy control, and evaluation aligned with clinical standards in m…