1 paper
Simin Chen, Pranav Pusarla, Baishakhi Ray
The rapid evolution of code largelanguage models underscores the need for effective and transparent benchmarking of their reasoning capabilities. However, the current benchmarking…