1 citations · 1 across the 4 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Detecting Data Contamination from Reinforcement Learning Post-training for Large Language Models
Yongding Tao, Tian Wang, Yihong Dong +4
Data contamination poses a significant threat to the reliable evaluation of Large Language Models (LLMs). This issue arises when benchmark samples may inadvertently appear in train…
cs.CL2026★ 1 cited
HumanLLM: Towards Personalized Understanding and Simulation of Human Nature
Yuxuan Lei, Tianfu Wang, Jianxun Lian +3
Motivated by the remarkable progress of large language models (LLMs) in objective tasks like mathematics and coding, there is growing interest in their potential to simulate human…
cs.CL2025
STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
Haiquan Hu, Jiazhi Jiang, Shiyou Xu +2
Evaluating large language models (LLMs) has become increasingly challenging as model capabilities advance rapidly. While recent models often achieve higher scores on standard bench…