1 paper
Zheyuan Yang, Zexi Kuang, Xue Xia +1
We introduce TestCase-Eval, a new benchmark for systematic evaluation of LLMs in test-case generation. TestCase-Eval includes 500 algorithm problems and 100,000 human-crafted solut…