6 papers
On the Robustness of Knowledge Editing for Detoxification
Ming Dong, Shiyi Tang, Ziyan Peng +2
Knowledge-Editing-based (KE-based) detoxification has emerged as a promising approach for mitigating harmful behaviours in Large Language Models. Existing evaluations, however, lar…
How Do People Quantify Naturally: Evidence from Mandarin Picture Description
Yayun Zhang, Guanyi Chen, Fahime Same +2
Quantification is a fundamental component of everyday language use, yet little is known about how speakers decide whether and how to quantify in naturalistic production. We investi…
Do Large Language Models Judge Error Severity Like Humans?
Diege Sun, Guanyi Chen, Zhao Fan +2
Large Language Models (LLMs) are increasingly used as automated evaluators in natural language generation, yet it remains unclear whether they can accurately replicate human judgme…
Emotional Supporters often Use Multiple Strategies in a Single Turn
Xin Bai, Guanyi Chen, Tingting He +2
Emotional Support Conversations (ESC) are crucial for providing empathy, validation, and actionable guidance to individuals in distress. However, existing definitions of the ESC ta…
CCNU at SemEval-2025 Task 3: Leveraging Internal and External Knowledge of Large Language Models for Multilingual Hallucination Annotation
Xu Liu, Guanyi Chen
We present the system developed by the Central China Normal University (CCNU) team for the Mu-SHROOM shared task, which focuses on identifying hallucinations in question-answering…
GPTEval: A Survey on Assessments of ChatGPT and GPT-4
Rui Mao, Guanyi Chen, Xulang Zhang +2
The emergence of ChatGPT has generated much speculation in the press about its potential to disrupt social and economic systems. Its astonishing language ability has aroused strong…