1 paper
Lilin Wang, Lucas Ramalho, Alan Celestino +6
Benchmarks like SWE-bench have standardized the evaluation of Large Language Models (LLMs) on repository-level software engineering tasks. However, these efforts remain limited by…