1 paper
Yifan Bai, Xiaoyang Liu, Zihao Mou +7
As large language models (LLMs) are increasingly deployed for software engineering, constructing high-quality benchmarks is crucial for evaluating not just the functional correctne…