1 paper · 1 filter
Wenjun Peng, Xinyu Wang, Qi Wu
Large language models (LLMs) have revolutionized automated code generation, yet the evaluation of their real-world effectiveness remains limited by static benchmarks and simplistic…