2 papers
cs.SE2025
On the need to perform comprehensive evaluations of automated program repair benchmarks: Sorald case study
Sumudu Liyanage, Sherlock A. Licorish, Markus Wagner +1
In supporting the development of high-quality software, especially necessary in the era of LLMs, automated program repair (APR) tools aim to improve code quality by automatically a…
cs.SE2025
Static Analysis as a Feedback Loop: Enhancing LLM-Generated Code Beyond Correctness
Scott Blyth, Sherlock A. Licorish, Christoph Treude +1
Large language models (LLMs) have demonstrated impressive capabilities in code generation, achieving high scores on benchmarks such as HumanEval and MBPP. However, these benchmarks…