1 paper
Elizabeth Dinella, Satish Chandra, Petros Maniatis
Large Language Models have demonstrated exceptional proficiency on coding tasks, but it is challenging to precisely evaluate their code reasoning ability. Existing benchmarks are i…