3 papers
cs.CL2026
CASTLE: A Comprehensive Benchmark for Evaluating Student-Tailored Personalized Safety in Large Language Models
Rui Jia, Ruiyi Lan, Fengrui Liu +7
Large language models (LLMs) have advanced the development of personalized learning in education. However, their inherent generation mechanisms often produce homogeneous responses…
cs.CL2026
Towards Comprehensive Stage-wise Benchmarking of Large Language Models in Fact-Checking
Hongzhan Lin, Zixin Chen, Zhiqi Shen +5
Large Language Models (LLMs) are increasingly deployed in real-world fact-checking systems, yet existing evaluations focus predominantly on claim verification and overlook the broa…
cs.CL2025
Is LLMs Hallucination Usable? LLM-based Negative Reasoning for Fake News Detection
Chaowei Zhang, Zongling Feng, Zewei Zhang +3
The questionable responses caused by knowledge hallucination may lead to LLMs' unstable ability in decision-making. However, it has never been investigated whether the LLMs' halluc…