4 papers
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models
Qianhong Guo, Wei Xie, Xiaofang Cai +7
Although large language models (LLMs) have shown exceptional capabilities across a wide range of tasks, reliable evaluation remains a critical challenge due to data contamination,…
AIPsychoBench: Understanding the Psychometric Differences between LLMs and Humans
Wei Xie, Shuoyoucheng Ma, Zhenhua Wang +4
Large Language Models (LLMs) with hundreds of billions of parameters have exhibited human-like intelligence by learning from vast amounts of internet-scale data. However, the unint…
Do Large Language Models Truly Grasp Mathematics? An Empirical Exploration From Cognitive Psychology
Wei Xie, Shuoyoucheng Ma, Zhenhua Wang +4
The cognitive mechanism by which Large Language Models (LLMs) solve mathematical problems remains a widely debated and unresolved issue. Currently, there is little interpretable ex…
HyperGo: Probability-based Directed Hybrid Fuzzing
Peihong Lin, Pengfei Wang, Xu Zhou +3
Directed grey-box fuzzing (DGF) is a target-guided fuzzing intended for testing specific targets (e.g., the potential buggy code). Despite numerous techniques proposed to enhance d…