14 citations · 22 across the 4 of their papers we have counts for
1 paper · 1 filter
Jialun Cao, Zhiyong Chen, Jiarong Wu +2
Code generation benchmarks such as HumanEval are widely adopted to evaluate LLMs' capabilities. However, after consolidating the latest 24 benchmarks, we noticed three significant…