activity
20242026
most citedBenchmark Data Contamination of Large Language Models: A Survey

20 citations · 20 across the 2 of their papers we have counts for

collaborators

5 papers

cs.AI2026

Phantom Gains: Auditing Self-Improvement Against a Measured Null

Cheng Xu, Nan Yan, Liming Chen +1

Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses. Tracking these transitions means diffe…

cs.CL2026

LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection

Cheng Xu, Changhong Jin, Yingjie Niu +5

The rapid development of Large Language Models (LLMs) has transformed fake news detection and fact-checking tasks from simple classification to complex reasoning. However, evaluati…

cs.CR2025

Cloud Digital Forensic Readiness: An Open Source Approach to Law Enforcement Request Management

Abdellah Akilal, M-Tahar Kechadi

Cloud Forensics presents a multi-jurisdictional challenge that may undermines the success of digital forensic investigations (DFIs). The growing volumes of domiciled and foreign la…

cs.CL2025

DCR: Quantifying Data Contamination in LLMs Evaluation

Cheng Xu, Nan Yan, Shuhao Guan +4

The rapid advancement of large language models (LLMs) has heightened concerns about benchmark data contamination (BDC), where models inadvertently memorize evaluation data during t…

cs.CL202420 cited

Benchmark Data Contamination of Large Language Models: A Survey

Cheng Xu, Shuhao Guan, Derek Greene +1

The rapid development of Large Language Models (LLMs) like GPT-4, Claude-3, and Gemini has transformed the field of natural language processing. However, it has also resulted in a…