1 paper
Ryan Deng, Yuanzhe Liu, Bastian Lipka +4
Large language model (LLM) agents now perform well on correctness-oriented repository-level tasks, including SWE-Bench issue resolution and feature implementation in real codebases…