Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models
Qianhong Guo, Wei Xie, Xiaofang Cai +7
Although large language models (LLMs) have shown exceptional capabilities across a wide range of tasks, reliable evaluation remains a critical challenge due to data contamination,…
cs.AI2025
Do Large Language Models Truly Grasp Mathematics? An Empirical Exploration From Cognitive Psychology
Wei Xie, Shuoyoucheng Ma, Zhenhua Wang +4
The cognitive mechanism by which Large Language Models (LLMs) solve mathematical problems remains a widely debated and unresolved issue. Currently, there is little interpretable ex…