1 paper
Tianze Han, Beining Xu, Hanbo Zhang +1
Mental-health dialogue models are increasingly evaluated by AI-based evaluators, yet these evaluators often treat surface empathy, supportiveness, or fluency as evidence of safety.…