1 paper
Miltiadis Allamanis, Sheena Panthaplackel, Pengcheng Yin
To evaluate code large language models (LLMs), research has relied on a few small manually curated benchmarks, such as HumanEval and MBPP, which represent a narrow part of the real…